Related Experiment Video
Updated: Jan 7, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Clinical Manifestations
Jet Mj Vonk1, Jiachen Lian2, Zoe Ezzes1
1University of California San Francisco (UCSF), San Francisco, CA, USA.
Background:
The non-fluent and logopenic variants of Primary Progressive Aphasia (nfvPPA, lvPPA;Figure 1) can be difficult to distinguish in early disease stages when expert clinicians need to correctly identify the specific nature of speech errors. NfvPPA is characterized by motor speech impairments and phonetic-motoric errors, while lvPPA involves phonological errors like sound substitutions and transpositions. Automated speech recognition (ASR) systems are promising tools to aid clinicians in objectively analyzing speech and language in dementia. However, current tools are inadequate in identifying specific speech errors, particularly at the phoneme level, as they prioritize fluent transcription omitting dysfluencies. We hypothesized that our novel forced alignment-based Scalable Speech Dysfluency Modeling Lightweight (SSDM-L) overcomes these limitations by capturing phoneme- and word-level disruptions, offering a data-driven alternative to perceptual judgment.
Method:
We analyzed recordings of reading aloud the Grandfather passage from 31 individuals with nfvPPA, 67 lvPPA, and 26 controls. Ten dysfluency variables were extracted using SSDM-L, including phoneme- and word-level insertions, replacements, repetitions, deletions, phoneme prolongations, and pauses between words. Unlike conventional ASR, SSDM-L incorporates a robust dysfluent- and time-aware phoneme and word transcriber, phonetic-aware subsequence aligner, and rule-based dysfluency detector to accurately identify phonemic errors with high precision. Group differences were assessed via MANCOVA, adjusting for age, education, and disease severity (MMSE, CDR-sum-of-boxes), and multiple comparisons. Backward stepwise logistic regression with repeated train-test splits identified robust predictors, and 5-fold cross-validation evaluated model performance.
Result:
All features except word repetition distinguished controls from PPA groups (p < .001-.012). NfvPPA and lvPPA differed in phoneme replacement (p = .036), phoneme deletions (p < .001), word replacement (p = .012), and word deletions (p < .001; Figure 2). These four features were also together selected as key predictors in logistic regression, yielding 81.6% accuracy (AUC = .761). Adding covariates yielded 80.5% accuracy (AUC = .843), while a covariates-only model achieved 65.2% accuracy (AUC = .659).
Conclusion:
Differentiating nfvPPA and lvPPA relies on detailed phoneme-level error characterization, traditionally requiring extensive expertise and manual effort. Our novel approach automates this process using a quick and easy-to-administer reading task, providing an objective, scalable, and clinically viable solution. By capturing dysfluencies missed by ASR, this method enhances diagnostic precision and reduces reliance on highly specialized clinicians, facilitating broader clinical adoption for differential diagnosis.
Related Concept Videos
Chronic Kidney Disease II: Clinical Manifestations
Coronary Artery Disease III: Clinical Manifestations
Endocarditis II: Clinical Features of Infective Endocarditis
Heart Failure III: Clinical Manifestations
Gastroesophageal Reflux Disease II: Clinical Features and Management
Clinical Manifestations
GERD presents itself in a multitude of ways, with symptoms varying from person to person. The hallmark symptoms are...
Hypertension III: Clinical Manifestations and Diagnostic Studies

