Related Experiment Video
Updated: Aug 5, 2026

08:32
Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
Published on: September 5, 2019
Multidimensional Prosodic and Semantic Coherence Modeling for Mandarin Mild Cognitive Impairment Detection
Rongyu Li1, Meihong Wu1,2
1School of Informatics, Xiamen University, 422 Siming South Road, Xiamen 361005, China.
Bioengineering (Basel, Switzerland)
|July 28, 2026
Summary
Early detection of Alzheimer's disease (AD) and mild cognitive impairment (MCI) is crucial. A new speech analysis model, Multi-Spec MCI-Net, shows promise for scalable, non-invasive screening by integrating multiple speech features.
Area of Science:
- Computational linguistics
- Artificial intelligence in healthcare
- Neuroscience
Background:
- Early detection of Alzheimer's disease (AD) and mild cognitive impairment (MCI) is vital for timely intervention.
- Current diagnostic methods like neuroimaging are expensive, invasive, and not scalable for population screening.
- Speech analysis offers a non-invasive, cost-effective alternative biomarker for cognitive assessment.
Purpose of the Study:
- To develop and evaluate a multimodal speech-based framework, Multi-Spec MCI-Net, for classifying healthy controls (HC) and individuals with MCI.
- To integrate diverse speech representations—token-level semantics, prosodic dynamics, and discourse coherence—for improved diagnostic accuracy.
- To create a clinically interpretable and scalable tool for early MCI screening.
Main Methods:
- Utilized a multimodal deep learning framework (Multi-Spec MCI-Net) combining dVAE, BERT, 1D-CNN with attention, and graph convolutional networks.
- Employed a gated fusion mechanism to adaptively weight three complementary speech modalities: semantics, prosody, and discourse coherence.
- Evaluated the model on the Chinese NCMMSC2021_AD dataset and the DementiaBank Mandarin subset, encompassing spontaneous dialogue and picture description tasks.
Main Results:
- Achieved 89.29% accuracy and 0.9584 ROC AUC on the NCMMSC2021_AD dataset, with 92.31% recall for MCI detection.
- Demonstrated robustness across different speech tasks, yielding 77.46% accuracy and 0.8280 AUC on a combined dataset.
- Multimodal fusion outperformed the semantic-only baseline by 5.16%, highlighting the contribution of each speech feature modality.
Conclusions:
- Multi-Spec MCI-Net provides an effective and interpretable approach for early MCI screening using speech analysis.
- The multimodal framework successfully integrates diverse linguistic dimensions for enhanced diagnostic performance.
- This speech-based method offers a scalable and non-invasive alternative for population-level cognitive impairment detection.

