Related Experiment Video
Updated: Sep 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Multimodal Dementia Prediction With Large Language Models: Cross-Attention Over Text, Audio, and Image
Felix Agbavor1, Hualou Liang1,2,3,4
1School of Biomedical Engineering and Science, Drexel University, Philadelphia, PA, United States.
Background:
Alzheimer disease (AD) is a leading cause of dementia, and there is growing interest in scalable approaches for early screening using speech-based tasks. While prior work has demonstrated promising results using either transcript-based language features or acoustic cues, most approaches remain unimodal or rely on simple fusion strategies that do not explicitly consider interactions across modalities.
Objective:
In this study, we propose an attention-based trimodal fusion framework that integrates text, audio, and image representations of the Cookie Theft picture, which serves as the shared visual stimulus in the picture-description task.
Methods:
Our method uses a new bidirectional cross-attention mechanism to achieve a unified multimodal embedding for downstream tasks. We evaluate the approach on 2 tasks: AD detection by classifying whether the participant has AD or not, and AD severity assessment by predicting Mini-Mental Status Examination cognitive scores.
Results:
On the AD detection task, trimodal fusion achieves the best overall performance (F1-score=0.8667, area under the receiver operating characteristic curve=0.9032), outperforming unimodal baselines, bimodal fusion, and conventional early or late fusion methods. For AD severity assessment, the proposed multimodal representation reduces prediction error of root mean squared error to about 4.20, improving over both unimodal and bimodal fusion settings. We further perform the ablation analysis to show that bidirectional cross-attention consistently outperforms conventional unidirectional cross-attention.
Conclusions:
These results demonstrate that attention-based multimodal fusion can enhance dementia prediction from picture-description responses and provide a strong foundation for developing multimodal cognitive screening pipelines.