Related Experiment Video
Updated: Aug 24, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Automatic speech recognition underperforms with diverse accents as measured by Word Error Rate and a semantic
Andrea Urqueta Alfaro1,2, Karina A Roundtree1, Mark S Pfaff1
1The MITRE Corporation, McLean, VA, USA.
Purpose:
Automated Speech Recognition (ASR) is used in assistive technologies such as caption phone and video calls for people who are deaf or hard of hearing. Evidence indicates that ASR's poor caption accuracy with "accented" speech (not the "standard" United States Broadcasting Mid-Western accent) is a barrier to successful communication. Efforts to improve ASR shortcomings should evaluate how caption errors alter the meaning intended by the speaker. Prior studies measured caption accuracy only using Word Error Rate (WER) which does not account for impact on meaning. This study expands existing work by assessing three major ASR engines' (Microsoft Azure, IBM Watson, Google Speech to Text) accuracy with a wide range of diverse accents using a measure of semantic similarity (SSM) between what the speaker said and how it was captioned, in addition to WER.
Results:
ASR caption accuracy is worse for accented speech compared to the Broadcasting Mid-Western accent, as measured by SSM and WER (p < 0.001). However, results were mixed when investigating whether some ASR engines perform better than others with accented speech; Azure performed better based on WER (p = .026), but not based on SSM.
Conclusion:
Findings support the need to include diverse accents in ASR engine training and assessment as well as evaluation of assistive technologies that leverage ASR applications. Results also support the use of the SSM in future research with accented speech, including investigating which accuracy measures are more appropriate (i.e., better alignment with human judgement) for comparing performance amongst ASR engines.
Related Concept Videos
Causes of Similarity-Dissimilarity Effect
Detection of Gross Error: The Q Test
