Related Experiment Video
Updated: Jul 17, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Cross-model disagreement as a reference-free signal for prioritizing human review in medical speech transcription
Abdolamir Karbalaie1, Fernando Seoane1,2,3,4, Farhad Abtahi1,2,5
1Department of Clinical Science, Intervention and Technology, Karolinska Institutet, Stockholm, Sweden.
Introduction:
Ambient AI scribes generate transcripts at scale, but routine quality assurance is constrained by the absence of human-verified reference transcripts in most deployment settings. We evaluated whether disagreement among heterogeneous automatic speech recognition (ASR) systems can serve as an informative signal for localizing transcription uncertainty, using a public English-language medical-speech corpus rather than clinical encounter recordings.
Methods:
Eight commercial and open-source ASR systems were applied to 50 medical-education audio clips (8 h 14 min). Multi-model outputs were aligned, and a leave-one-out consensus procedure was used to score per-model agreement while reducing circularity.
Results:
Disagreement across models was sparse and localized: 72.1% of positions showed strong agreement (7-8 systems concordant), whereas only 2.5% were high-risk positions with minimal agreement (0-3 systems). Low-agreement regions were systematically enriched for meaning-bearing lexical differences, defined as lexical mismatches after excluding punctuation, contraction, numeric, and filler variation. A single-annotator human-corrected (HC) validation layer showed that transcription errors increased monotonically with decreasing agreement. At an illustrative post hoc threshold, flagging positions where six or fewer systems agreed selected 28.6% of tokens while recovering 93.7% of single-annotator HC-verified errors on this proxy corpus.
Discussion:
These findings suggest that cross-model disagreement may help focus human review on a small number of likely error-prone transcript regions. However, agreement among all systems does not guarantee correctness, because shared errors may remain undetected by this approach. Validation on real clinical encounter data is required before operational deployment.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy