Related Experiment Video
Updated: Jul 1, 2025

A Protocol for Comprehensive Assessment of Bulbar Dysfunction in Amyotrophic Lateral Sclerosis ALS
Published on: February 21, 2011
Evaluating Fluency in Aphasia: Fluency Scales, Trichotomous Judgements, or Machine Learning
Jeet Metu1, Vishal Kotha2, Argye E Hillis3
1Rock Ridge High School, Johns Hopkins University School of Medicine, and Cognitive Science, Johns Hopkins University, Baltimore, MD 21287.
A simple human judgment of fluent, non-fluent, or mixed speech in aphasia is more reliable than the Western Aphasia Battery-Revised (WAB-R) fluency scale or machine learning models. Further research is needed before machine learning can reliably assess spoken language fluency.
Area of Science:
- Neurolinguistics
- Computational linguistics
- Speech-language pathology
Background:
- Aphasia classification relies on tools like the Western Aphasia Battery-Revised (WAB-R).
- The WAB-R's fluency scale lacks objectivity and has low inter-rater reliability due to subjective weighting of speech dimensions.
- This subjectivity impacts accurate aphasia classification and may be addressed by machine learning.
Purpose of the Study:
- To investigate if convolutional and recurrent neural networks can reliably classify fluent and non-fluent aphasia.
- To compare the reliability of machine learning models against the WAB-R fluency scale and expert clinician judgment.
Main Methods:
- Developed and trained convolutional and recurrent neural network models on public domain speech samples.
- Validated models using speech data from post-stroke aphasia participants.
- Assessed inter-rater reliability using Kappa scores among speech-language pathologists (SLPs) and between SLPs and the models.
Main Results:
- Machine learning models achieved high accuracy in detecting fluent (83%) and non-fluent (81%) speech on test datasets.
- Inter-rater reliability among SLPs using the WAB-R scale was poor to perfect for specific scores but substantial for broad categories (fluent vs. non-fluent).
- Agreement between models and SLPs was moderate for non-fluent and substantial for fluent speech; however, SLP trichotomous judgment (fluent, non-fluent, mixed) showed almost perfect agreement.
Conclusions:
- Simple trichotomous judgment by SLPs (fluent, non-fluent, mixed) demonstrated higher reliability and validity than the WAB-R fluency scale and machine learning models.
- The WAB-R fluency scale's utility for aphasia classification warrants reconsideration.
- Current machine learning approaches are not yet reliable enough for rating spoken language fluency in aphasia.
More Related Videos
10:15Utilizing Repetitive Transcranial Magnetic Stimulation to Improve Language Function in Stroke Patients with Chronic Non-fluent Aphasia
Published on: July 2, 2013
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024