Related Experiment Video
Updated: Jan 7, 2026

Transcranial Direct Current Stimulation tDCS for Memory Enhancement
Published on: September 18, 2021
Dementia Care Research and Psychosocial Factors
Yexuan Xiao1, Qianhui Pan1, Nan Jiang1
1Tsinghua University School of Healthcare Management, Beijing, Beijing, China.
Background:
The demand for accessible and actionable information to support mild cognitive impairment (MCI) management is growing, particularly for care partners and non-psychiatric healthcare professionals seeking reliable guidance. Large Language Models (LLMs) have shown potential across various domains but have not been systematically evaluated in addressing MCI-related queries. A critical gap exists in understanding LLM performance across different aspects of dementia care and the impact of language differences on their effectiveness. This study addresses these gaps by assessing LLM-generated responses in both English and Chinese, providing essential insights into their capabilities and limitations.
Method:
Healthcare professionals (N = 5) and care partners (N = 5) were recruited from diverse clinical and caregiving backgrounds. A set of 72 open-ended questions across four MCI domains-Symptoms, Treatment, Care partner support, and Nursing&rehabilitation-was developed. Responses from three LLMs (ChatGPT-4, Gemini, Kimi) were evaluated by a five-point accuracy scale using a double-blind design. Ratings were based on accuracy, comprehensibility, specificity, and actionability. Statistical analyses included ICCs for reliability and Mann-Whitney U tests to compare responses.
Result:
Among the four domains of MCI management, ChatGPT had the best performance in symptoms-related queries (average score: 4.03, all p < 0.001). All three LLM-generated responses demonstrated better alignment with healthcare professionals' requirements, achieving significantly higher concordance in comprehensibility (HP:4.24 vs CP:4.06, p <0.001) and actionability (HP:4.04 vs CP:3.72, p <0.001). Performance parity was observed in accuracy (CP:4.32 vs HP:4.27) and specificity (CP:3.79 vs HP:3.83) with no statistical significance. Additionally, English responses surpassed Chinese responses in accuracy (4.32 vs. 4.26, p = 0.056), comprehensibility (4.23 vs. 4.10, p < 0.001), and specificity (3.96 vs. 3.66, p < 0.001), while actionability scores showed no significant difference (3.94 vs. 3.89, p = 0.239).
Conclusion:
This study demonstrates LLMs' proficiency in symptom-related inquiries and stronger alignment with healthcare professionals' operational needs versus care partners' accuracy priorities. English responses outperformed Chinese due to corpus disparities, highlighting needs for language-specific optimization. Findings underscore LLMs' clinical potential, urging enriched Chinese medical corpora and specialized model development to address care partners' unmet requirements.
More Related Videos
08:36The Immersive Cleveland Clinic Virtual Reality Shopping Platform for the Assessment of Instrumental Activities of Daily Living
Published on: July 28, 2022
10:13Assessment of Age-related Changes in Cognitive Functions Using EmoCogMeter, a Novel Tablet-computer Based Approach
Published on: February 14, 2014
Related Concept Videos
Dementia
The progression of dementia is generally gradual....
Psychological and Sociocultural Causes of Schizophrenia
Alzheimer's Disease: Overview
The clinical diagnosis of AD hinges on the presence of memory and other cognitive impairments. Biomarkers, such as changes in Aβ...
Alzheimer's Disease: Treatment
Cognitive Development During Adulthood
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities