Related Experiment Video
Updated: Jun 11, 2025

12:18
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
7.5K
A multimodal cross-transformer-based model to predict mild cognitive impairment using speech, language and vision.
Farida Far Poor1, Hiroko H Dodge2, Mohammad H Mahoor1
1Department of Electrical and Computer Engineering, University of Denver, Denver, CO, USA.
Computers in Biology and Medicine
|September 27, 2024
Summary
This study introduces a novel AI approach using multimodal data to accurately detect Mild Cognitive Impairment (MCI), an early stage of cognitive decline. The method significantly improves prediction accuracy compared to single-data approaches, aiding early diagnosis.
Area of Science:
- Artificial Intelligence
- Neuroscience
- Gerontology
Background:
- Mild Cognitive Impairment (MCI) is a transitional stage between normal cognition and dementia.
- Early detection of MCI is crucial to prevent progression to Alzheimer's disease and other dementias.
- Existing AI methods for MCI detection often rely on unimodal data, limiting prediction accuracy.
Purpose of the Study:
- To develop and evaluate a robust multimodal AI architecture for accurate MCI prediction.
- To address the challenge of effectively fusing information from different data modalities for enhanced detection.
- To improve upon existing unimodal and bimodal approaches for MCI identification.
Main Methods:
- Proposed a Deep Learning-based method integrating speech, language, and vision modalities.
- Utilized an embedding-level fusion architecture with a co-attention mechanism for inter-modal relationship preservation.
- Employed the I-CONECT dataset comprising semi-structured conversations from individuals aged 75+.
Main Results:
- The multimodal fusion model achieved an average AUC of 85.3% in differentiating MCI from Normal Cognition (NC).
- This significantly outperformed unimodal models (60.9% AUC) and bimodal models (76.3% AUC).
- The co-attention mechanism effectively captured complementary information across speech, language, and vision data.
Conclusions:
- The proposed multimodal fusion approach offers a more accurate and reliable method for early MCI detection.
- Leveraging combined speech, language, and vision data significantly enhances predictive performance.
- This AI-driven strategy holds promise for clinical applications in identifying individuals at risk of cognitive decline.

