Related Experiment Video
Updated: Jan 9, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
TransformerCARE: A novel speech analysis pipeline using transformer-based models and audio augmentation techniques
Hossein Azadmaleki1, Ali Zolnour1, Sina Rashidi1
1Columbia University Irving Medical Center, 622 W 168th St, New York, NY 10032, United States.
Objective:
Early diagnosis of cognitive impairment, including Alzheimer's and other dementias, is critical for effective treatment and slowing disease progression. However, over 50% of cases remain undiagnosed until advanced stages due to limitations in current methods. Recognizing speech impairments as early markers of cognitive decline, this study evaluated the utility of speech analysis as a technique for early detection. We introduce TransformerCARE, a speech processing pipeline utilizing advanced speech transformer models.
Methods:
TransformerCARE incorporated a series of key steps, including preprocessing, speech segmentation, transformer fine-tuning, segment aggregation, performance evaluation, and data augmentation. In the fine-tuning step, we evaluated the performance of four state-of-the-art speech transformer models: Wav2vec 2.0, HuBERT, WavLM, and DistilHuBERT. For data augmentation, we adopted multiple techniques, with particular emphasis on frequency masking due to its ability to preserve subtle acoustic cues associated with cognitive impairment. We measured the performance of TransformerCARE on the ADReSSo Challenge dataset from DementiaBank, comprising 237 subjects (122 cognitively impaired and 115 cognitively normal).
Results:
TransformerCARE demonstrated its highest performance with HuBERT, achieving an AUC of 81.80 (F1-score = 79.31) using an aggregation technique that averaged embeddings of 14-second speech segments. Augmenting the training data with frequency masking improved performance by 5 %, resulting in an AUC of 86.11 (F1-score = 84.63). We also demonstrated that incorporating clinicians' speech during patient interactions can improve the performance of the pipeline. Our error analysis revealed significant differences between the acoustic patterns of correctly identified negative cases (true negatives) and those incorrectly identified as positive (false positives), as well as between correctly identified positive cases (true positives) and those incorrectly identified as negative (false negatives). This indicates specific deviations in speech characteristics among inaccurately diagnosed subjects.
Conclusion:
In summary, TransformerCARE demonstrates strong potential for integration into clinical workflows as a screening tool for cognitive impairment, aiding in the timely and appropriate care of affected patients.
More Related Videos
09:47Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025