Benchmarking Automatic Speech Recognition Technology for Natural Language Samples of Children With and Without
Summary
Automatic speech recognition (ASR) shows promise for natural language sampling (NLS) research but struggles with young children, especially those with Down syndrome (DS). Human oversight is crucial for accurate clinical transcription of diverse speech patterns.
Area of Science:
- Clinical Linguistics
- Speech Technology
- Developmental Psychology
Background:
- Natural Language Sampling (NLS) provides valuable real-world speech data but relies on costly human transcription.
- Automatic Speech Recognition (ASR) offers a potential solution to streamline NLS data processing.
- The efficacy of ASR in clinical settings, particularly with young children and those with developmental delays, is largely unexamined.
Purpose of the Study:
- To evaluate the performance of the OpenAI Whisper ASR model for transcribing speech data from toddlers.
- To compare ASR accuracy in children with Down syndrome (DS) versus typically developing (TD) children.
- To assess ASR capabilities in capturing developmentally relevant speech features beyond words.
Main Methods:
- Thirty-four Natural Language Sampling (NLS) sessions were analyzed.
- Participants included toddlers with Down syndrome (DS; n=19; 2-5 years) and typically developing (TD) toddlers (n=15; 2-3 years).
- OpenAI Whisper ASR outputs were manually compared against human transcriptions.
Main Results:
- ASR accurately transcribed 50% of words for TD children but only 14% for children with DS.
- Word error rates included approximately 20% missed words and 21% replaced words (TD) vs. 6% (DS).
- ASR failed to capture nearly 50% of non-speech vocalizations in the DS group and misinterpreted most others.
Conclusions:
- While ASR can reduce transcription time, its current limitations necessitate human-in-the-loop systems for clinical research.
- ASR performance is significantly lower for children with Down syndrome, highlighting disparities in underrepresented groups.
- Further development of ASR is needed to accurately process diverse and developmentally informative speech patterns in clinical populations.


