Related Experiment Video
Updated: Sep 5, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
A Cautious Integration With AI in the Clinic: A Standardized-Patient Pilot Study of ChatGPT's Reliability in Hamilton
Chun-Hung Chang1,2,3, Szu-Wei Cheng4, Wei-Jen Chen1
1Department of Psychiatry, An-Nan Hospital, China Medical University, 709 Tainan, Taiwan.
Background:
Artificial intelligence (AI) integration offers significant potential to improve mental healthcare, however, the reliability of large language models (LLMs) in performing nuanced clinical tasks remains an important and largely unanswered question. This study aimed to evaluate ChatGPT's performance in scoring the Hamilton Depression Rating Scale (HAMD-21) compared with expert raters using standardized patients (SPs).
Methods:
Three senior mental health experts created and portrayed scenarios for ten SPs representing diverse depressive symptom profiles. Recorded interviews were transcribed and used as input for ChatGPT-4o. HAMD-21 scores generated by ChatGPT were compared with those assigned by expert raters and with predefined script-based reference scores. Inter-rater reliability was assessed using intraclass correlation coefficient (ICC), and differences between raters were evaluated using Steiger's tests.
Results:
ChatGPT and the expert raters achieved good-to-excellent reliability for total HAMD-21 scores (experts: ICC = 0.9921; ChatGPT: ICC = 0.9739). However, expert raters achieved perfect ICCs on 11 individual items, whereas ChatGPT achieved perfect agreement on only 2 items. Steiger's test demonstrated that experts significantly outperformed ChatGPT on 10 individual items as well as on total scores (Z = 1.931, p = 0.0268). Qualitative review revealed that ChatGPT tended to overestimate scores on items related to insomnia and somatic symptoms (items 4-6 and 13) and frequently miscalculated total scores.
Conclusions:
ChatGPT demonstrated excellent agreement on total HAMD-21 scores in structured, text-based depression assessments, supporting the potential role of LLMs as adjunctive tools for standardized depression severity evaluation. However, item-level discrepancies and systematic scoring errors indicate that human oversight remains essential for clinically nuanced interpretation.
More Related Videos
07:14Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
13:18Robotically Delivered fMRI-Guided Personalized Transcranial Magnetic Stimulation Therapy for Treatment-Resistant Depression
Published on: April 10, 2026