Related Experiment Video
Updated: Jun 13, 2025

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis ALS
Published on: February 21, 2011
Automatic Speech Recognition in Primary Progressive Apraxia of Speech
Katerina A Tetzloff1, Daniela Wiepert1, Hugo Botha1
1Department of Neurology, Mayo Clinic, Rochester, MN.
Automatic speech recognition (ASR) effectively differentiates primary progressive apraxia of speech (PPAOS) from healthy speech and correlates with severity. However, ASR struggles to distinguish PPAOS subtypes and shows limited agreement with manual transcriptions.
Area of Science:
- Speech-language pathology
- Computational linguistics
- Neurology
Background:
- Diagnosing motor speech disorders like primary progressive apraxia of speech (PPAOS) involves transcribing speech, a labor-intensive process requiring expert listeners.
- Automatic Speech Recognition (ASR) systems offer a potential solution for efficient diagnosis and monitoring of PPAOS.
- Evaluating the efficacy of readily available ASR systems for PPAOS speech transcription is crucial.
Purpose of the Study:
- To assess the effectiveness of the wav2vec 2.0 ASR system in transcribing PPAOS speech.
- To determine if ASR-derived Word Error Rate (WER) can differentiate PPAOS patients from healthy controls and among PPAOS subtypes.
- To investigate the correlation between WER and PPAOS severity, and compare ASR errors to manual transcription errors.
Main Methods:
- Recorded 45 PPAOS patients and 22 healthy controls repeating 13 words.
- Transcribed recordings manually and using the wav2vec 2.0 ASR system.
- Compared WER, phonetic, and prosodic errors between groups and against manual transcriptions.
Main Results:
- Mean WER was significantly higher in PPAOS patients (0.88) than controls (0.33).
- WER correlated with PPAOS severity and distinguished patients from controls, but not PPAOS subtypes.
- ASR phonetic/prosodic errors did not differentiate subtypes, unlike human transcriptions, with poor agreement between ASR and human error counts.
Conclusions:
- ASR, specifically wav2vec 2.0, is valuable for distinguishing disordered from healthy speech and assessing PPAOS severity.
- ASR's current limitations include an inability to differentiate PPAOS subtypes and weak agreement with manual transcriptions.
- ASR can aid PPAOS speech transcription, but its use requires careful consideration of its limitations and the specific research questions.
More Related Videos
05:48Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
08:17A Semantic Priming Event-related Potential ERP Task to Study Lexico-semantic and Visuo-semantic Processing in Autism Spectrum Disorder
Published on: April 12, 2018