Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Exploring the impact of noise, language familiarity, and experimental settings on emotion recognition.

Frontiers in psychology·2025
Same author

Formant-based vowel categorization for cross-lingual phone recognition.

The Journal of the Acoustical Society of America·2025
Same author

Neural Representations of Non-native Speech Reflect Proficiency and Interference from Native Language Knowledge.

The Journal of neuroscience : the official journal of the Society for Neuroscience·2023
Same author

Recognizing non-native spoken words in background noise increases interference from the native language.

Psychonomic bulletin & review·2022
Same author

The time course of adaptation to distorted speech.

The Journal of the Acoustical Society of America·2022
Same author

The differential roles of lexical and sublexical processing during spoken-word recognition in clear and in noise.

Cortex; a journal devoted to the study of the nervous system and behavior·2022

Related Experiment Video

Updated: Apr 27, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

1.1K

Speech recognition performance disparities between Dutch diverse speaker groups.

Yuanyuan Zhang1, Thomas De Valck1, Odette Scharenborg1

  • 1Multimedia Computing Group, Delft University of Technology, Postbus 5031, 2600 GA, Delft, The Netherlands.

Phonetica
|April 26, 2026
PubMed
Summary

Automatic speech recognition (ASR) systems struggle with diverse Dutch speech. Performance disparities are mainly due to language proficiency and speech impairment, not demographics, and system updates don't always help.

Keywords:
Dutch diverse speechautomatic speech recognitiondysarthric speechnon-native accentsperformance disparities

More Related Videos

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K
An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
07:56

An Automated System for Sound Localization Testing in Hearing-Impaired Listeners

Published on: March 13, 2026

176

Related Experiment Videos

Last Updated: Apr 27, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

1.1K
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K
An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
07:56

An Automated System for Sound Localization Testing in Hearing-Impaired Listeners

Published on: March 13, 2026

176

Area of Science:

  • Speech Technology
  • Human-Computer Interaction
  • Computational Linguistics

Background:

  • State-of-the-art automatic speech recognition (ASR) systems excel with typical speech.
  • Performance degrades significantly for diverse speech, influenced by demographic and sociolinguistic factors.
  • Rapid advancements in ASR necessitate continuous evaluation on varied speech inputs.

Purpose of the Study:

  • To evaluate the performance of recent ASR systems and custom models on diverse Dutch speech.
  • To identify factors contributing to performance disparities among different speaker groups.
  • To assess the impact of data processing and system updates on ASR performance for diverse users.

Main Methods:

  • Tested nine commercial ASR systems (Google, Microsoft, Meta, NVIDIA, OpenAI) and three custom models.
  • Evaluated performance on Dutch diverse speech datasets.
  • Analyzed recognition results to pinpoint factors influencing performance variations across speaker groups.

Main Results:

  • All evaluated ASR systems exhibited similar patterns in performance disparities for diverse Dutch speakers.
  • Language proficiency and severe speech motor impairment had a greater impact than demographic factors.
  • Differences in data processing and decoding significantly affected recognition accuracy.
  • System updates did not consistently improve performance or reduce disparities for diverse groups.

Conclusions:

  • Acoustic variability from demographic/sociolinguistic factors appears adequately represented in typical training data.
  • Language proficiency and speech impairment are key challenges for ASR systems with diverse Dutch speakers.
  • Ongoing ASR development must prioritize equitable performance across all user groups, not just typical speech.