Related Experiment Video
Updated: Sep 28, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Artificial intelligence-based speech error detection to differentiate primary progressive aphasia variants
Jet M J Vonk1, Jiachen Lian2, G Lynn Kurteff1
1Edward and Pearl Fein Memory and Aging Center, Department of Neurology, University of California San Francisco, San Francisco , CA 94158, USA.
Abstract:
Artificial intelligence-based approaches to speech analysis have the potential to assist with objective speech error analysis in aphasia but off-the shelf tools often fail to detect speech errors due to prioritizing 'fluent transcription'. Speech production errors (dysfluencies) are hallmark diagnostic features of the nonfluent and logopenic variants of primary progressive aphasia, yet they can be challenging to detect and characterize even by expert clinicians. This study aimed to evaluate whether the novel automated lightweight Scalable Speech Dysfluency Modeling system, specifically designed to detect dysfluencies, could accurately distinguish primary progressive aphasia variants using voice recordings of individuals reading a brief passage. Participants included a total of 104 individuals, 40 with non-fluent primary progressive aphasia and 40 with logopenic primary progressive aphasia (matched on disease severity), and 24 healthy controls who read aloud the 'Grandfather Passage' as part of a widely used motor speech evaluation. We automatically extracted ten speech error (dysfluency) variables, including insertions, replacements, and deletions at both phoneme- and word-levels, and phoneme-level prolongations and repetitions. Group differences were assessed via ANOVAs controlling for age, education, and disease severity (Mini-Mental State Examination and Clinical Dementia Rating sum-of-boxes). To test clinical relevance, we performed correlation analyses with motor speech evaluation ratings provided by experienced speech-language pathologists (i.e. gold standard) within the non-fluent primary progressive aphasia group. Classification performance was assessed by training random forest and XGBoost machine-learning models including 5-fold cross-validation. All individuals read the entire passage in less than five minutes. The lightweight Scalable Speech Dysfluency Modeling system detected 8 of the 10 predefined dysfluency features at sufficient frequency to include them in subsequent analyses. All eight features distinguished primary progressive aphasia from controls (P < 0.006). Individuals with non-fluent primary progressive aphasia made more errors than logopenic primary progressive aphasia on every feature (all P < 0.023). Each feature showed a moderate positive correlation with a global motor speech evaluation apraxia/dysarthria score (r = 0.31-0.56; P < 0.001-0.053). Together, the eight features were able to classify non-fluent versus logopenic at area under the receiver operating characteristic curve = 0.798 (held-out random forest), 0.699 (held-out XGBoost), 0.714 (cross-validated random forest), and 0.704 (cross-validated XGBoost). In sum, automated speech error analysis accurately distinguished non-fluent and logopenic variants using a brief reading task. This quick error-sensitive scalable artificial intelligence system has the potential of providing a practical tool to aid diagnosis in aphasia and motor speech disorders.
