Related Experiment Video
Updated: Sep 12, 2025

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Multimodal Alzheimer's disease recognition from image, text and audio
Byounghwa Lee1, Hwa Jeon Song2, Young-Jin Park3
1Integrated Intelligence Research Section, Electronics and Telecommunications Research Institute, Daejeon, 34129, Republic of Korea. byounghwa.lee@etri.re.kr.
This study introduces a new multimodal model for Alzheimer's disease (AD) prediction, integrating image, text, and audio data. The novel approach significantly improves diagnostic accuracy by combining visual context with speech and text analysis.
Area of Science:
- Neuroscience
- Artificial Intelligence
- Medical Diagnostics
Background:
- Alzheimer's disease (AD) is a progressive neurodegenerative disorder impacting cognitive function.
- Current AD diagnosis often relies on analyzing verbal descriptions, primarily using speech and text-based models.
- Integrating visual context into AD diagnostic models remains an underexplored area.
Purpose of the Study:
- To propose and evaluate a novel multimodal model for Alzheimer's disease prediction.
- To integrate image, text, and audio data for enhanced diagnostic accuracy.
- To investigate the cooperative contribution of different modalities in AD classification.
Main Methods:
- Developed a multimodal model incorporating image, text, and audio data.
- Utilized a vision-language model for image and text processing, structured as a bipartite graph.
- Employed co-attention-based intermediate fusion and late fusion for integrating all three modalities.
- Conducted an ablation study using Shapley values to quantify modality contributions and develop an auxiliary loss function.
Main Results:
- The proposed multimodal model achieved a diagnostic accuracy of 90.61%, surpassing existing state-of-the-art methods.
- Shapley value analysis informed an auxiliary loss function that adaptively adjusted modality importance during training.
- Analysis of attention patterns revealed that audio and text provide complementary information for AD classification.
Conclusions:
- Integrating image, text, and audio modalities through co-attention-based fusion significantly enhances Alzheimer's disease classification performance.
- The multimodal approach offers a more comprehensive diagnostic tool compared to unimodal methods.
- Audio and text modalities offer complementary diagnostic cues, highlighting the value of multimodal data integration.
Related Concept Videos
Alzheimer's Disease: Overview
The clinical diagnosis of AD hinges on the presence of memory and other cognitive impairments. Biomarkers, such as changes in Aβ...
Alzheimer's Disease: Treatment
Dementia
The progression of dementia is generally gradual....

