Related Experiment Video
Updated: Aug 5, 2026

12:11
Objectively Assessing Sports Concussion Utilizing Visual Evoked Potentials
Published on: April 27, 2021
Speech-based concussion detection in athletes using Mel-spectrograms and convolutional neural networks
Rahmina Rubaiat1, Christian Poellabauer1
1Mobile Sensing and Analytics Lab, Knight Foundation School of Computing and Information Sciences, Florida International University, Miami, FL, United States.
Frontiers in Neurology
|July 31, 2026
Summary
Deep learning analysis of speech tasks shows promise for detecting concussion. Dynamic speech features, like reading and syllable repetition, achieved high accuracy in identifying neurological dysfunction.
Area of Science:
- Neurology
- Artificial Intelligence
- Digital Health
Background:
- Neurological disorders can cause subtle changes in motor control, cognition, and speech.
- Conventional diagnostic tools and imaging (CT, MRI) often miss these functional abnormalities.
- Non-invasive digital health approaches are gaining traction for neurological assessment.
Purpose of the Study:
- To investigate the efficacy of speech-based deep learning for detecting concussion-related dysfunction.
- To analyze structured speech tasks for identifying subtle neurological alterations.
- To explore the potential of convolutional neural networks (CNNs) in speech analysis for mild traumatic brain injury (concussion).
Main Methods:
- Audio recordings of 225 athletes performing eight structured speech tasks were collected.
- Recordings were converted to Mel spectrograms and analyzed using a CNN with data augmentation.
- Gradient-weighted class activation mapping (Grad-CAM) was used to interpret model behavior and task relevance.
Main Results:
- The CNN model demonstrated strong classification performance, with accuracy up to 96% for specific speech tasks.
- Tasks involving dynamic articulation and rhythmic control (e.g., sentence reading, rapid syllable repetition) yielded the highest accuracy.
- Sustained vowel tasks showed lower performance and less focused model attention due to limited acoustic variability.
Conclusions:
- CNN-based analysis of dynamic speech features is effective for detecting neurological dysfunction.
- Task selection is crucial for optimizing speech-based digital diagnostic tools.
- This framework shows potential for future speech-based neurological assessment, with concussion as an initial application.

