Related Experiment Video
Updated: Apr 24, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Evaluating large language models for mild cognitive impairment among older adults: A bilingual comparison of ChatGPT,
Yexuan Xiao1,2,3, Qianhui Pan2,3, Haoyuan Liu2,3
1Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University, Beijing, China.
Abstract:
Objective: To evaluate large language models (LLMs) in managing mild cognitive impairment (MCI) and supporting nonspecialist healthcare professionals and care partners, comparing English and Chinese responses. Methods: Seventy-two MCI-related questions were submitted to ChatGPT-4o, Gemini, and Kimi. Responses were assessed for accuracy, comprehensibility, specificity, and actionability using a 5-point Likert scale. Statistical analyses included intraclass correlation coefficients and Mann-Whitney U tests. Results: LLMs performed best in the symptoms and diagnosis domain (M = 4.11 ± 0.15). Healthcare professionals' needs were better met than those of care partners, particularly in comprehensibility and actionability (p < .001). English responses were significantly more comprehensible and specific than Chinese responses (p < .001). Conclusion: This study highlights the potential of LLMs like ChatGPT, Gemini, and Kimi in supporting MCI management, especially in diagnosis and providing actionable insights. However, their performance varied across languages and user groups, with English responses generally more effective than Chinese. The findings emphasize the need for culturally and linguistically adapted LLMs to enhance accuracy and usability. Future research should focus on expanding user diversity, improving adaptability, and incorporating region-specific data to optimize LLMs for MCI care.
More Related Videos
Related Concept Videos
Language and Cognition
Cognitive Development During Adulthood

