Related Experiment Video
Updated: Mar 28, 2026

05:19
Practical Methodology of Cognitive Tasks Within a Navigational Assessment
Published on: June 1, 2015
14.1K
Large Language Model Adaptation Strategies in Speech-Based Cognitive Screening: Systematic Evaluation.
Fatemeh Taherinezhad1, Mohamad Javad Momeni Nezhad1, Sepehr Karimi1
1Columbia University Irving Medical Center, 622 W, 168th St, New York, NY, 10032, United States.
JMIR AI
|March 26, 2026
Summary
Large language models effectively screen for Alzheimer disease and related dementias (ADRD) using speech. Token-level fine-tuning generally yields the best results for scalable, accurate detection.
Area of Science:
- Natural Language Processing
- Artificial Intelligence in Healthcare
- Speech Analysis
Background:
- Over 50% of US adults with Alzheimer disease and related dementias (ADRD) are undiagnosed.
- Speech-based screening offers a scalable solution, but the effectiveness of large language model (LLM) adaptation strategies requires further investigation.
Purpose of the Study:
- To compare various LLM adaptation strategies for detecting cognitive impairment using speech data.
- To evaluate both text-only and multimodal LLM approaches on DementiaBank datasets.
Main Methods:
- Analysis of audio-recorded speech from 237 participants (ADRD vs. cognitive normal) in the ADReSSo dataset.
- Evaluation of nine text-only LLMs and three multimodal models using adaptation strategies: in-context learning (ICL), reasoning-augmented prompting, and parameter-efficient fine-tuning.
- Assessment of strategy generalizability on the DementiaBank Delaware dataset (mild cognitive impairment vs. cognitive normal).
Main Results:
- Prototype demonstrations in ICL yielded the highest performance (F1-score up to 0.81) on the ADReSSo dataset.
- Token-level fine-tuning achieved the highest scores across models (e.g., LLaMA 3B: F1=0.83, AUC=0.91).
- Reasoning-augmented prompting benefited smaller models, while multimodal models did not outperform top text-only systems.
Conclusions:
- Detection accuracy depends on demonstration selection, reasoning design, and tuning methods.
- Token-level fine-tuning is generally most effective for speech-based ADRD and mild cognitive impairment screening.
- Adapted open-weight models show potential to match or surpass commercial LLMs; multimodal models may need further refinement.

