HiACC: Hinglish adult & children code-switched corpus
Shruti Singh1, Muskaan Singh2, Virender Kadyan1
1SoCS, University of Petroleum and Energy Studies, Dehradun, Uttarakhand, India.
This study introduces the HiACC corpus, a new Hinglish speech dataset for improving automatic speech recognition (ASR) systems. The corpus, featuring adult and child speakers, addresses a critical gap in resources for code-switched communication.
Area of Science:
- Computational Linguistics
- Speech Processing
- Sociolinguistics
Background:
- Code-switching, the alternation between languages, is common among bilinguals, particularly in India with Hinglish (Hindi-English).
- Existing automatic speech recognition (ASR) systems struggle with code-switched data due to a lack of diverse, representative datasets, leading to higher word error rates (WER).
- Current ASR models show a 30-50% increase in WER when processing code-switched speech compared to monolingual speech.
Purpose of the Study:
- To introduce the HiACC corpus, a benchmark Hinglish speech dataset designed to enhance ASR performance in resource-constrained environments.
- To address the scarcity of publicly available code-switched speech datasets, especially those including children's speech.
- To provide a valuable resource for linguistic and computational analysis of code-switching.
Main Methods:
- Development of the HiACC corpus, comprising 3,318 audio segments from adults and 1,858 from children.
- Inclusion of 5.24 hours of read and spontaneous Hinglish speech with detailed annotations and code-switching tags.
- Public release of the corpus with segmented audio and aligned transcripts for open research.
Main Results:
- Baseline ASR experiments demonstrated that models trained on monolingual data underperform significantly, with a WER of approximately 42% on the Hinglish test set.
- The HiACC corpus provides the first publicly available resource for code-switched Hinglish speech that includes both adult and child speakers.
- The dataset's availability is expected to catalyze progress in developing more robust ASR systems for code-switched languages.
Conclusions:
- The HiACC corpus is a crucial resource for advancing research in automatic speech recognition for code-switched languages like Hinglish.
- The findings highlight the challenges ASR systems face with code-switched input and the need for specialized datasets.
- This work paves the way for more inclusive and accurate speech recognition technologies for diverse linguistic communities.
More Related Videos
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
08:32Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
Published on: September 5, 2019
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Components of Language
Language and Cognition
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Genetic Lingo
Cognitive Development During Adulthood
