Related Experiment Video
Updated: Aug 11, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Measuring Depression Severity With Clinical Global Impression-Severity Scale Scores From Clinical Notes Using Large
Kevin Li1, Ayah Zirikly2,3, Sarah C Collica1
1Department of Psychiatry and Behavioral Sciences, School of Medicine, Johns Hopkins University, 550 N Broadway, Baltimore, MD, 21205, United States, 1 410-614-1923.
JMIR Formative Research
|August 10, 2026
Summary
Large language models (LLMs) can accurately estimate Clinical Global Impression-Severity (CGI-S) scores from psychiatric notes for major depressive disorder (MDD). GPT-4o demonstrated high agreement with expert ratings, suggesting potential for scalable outcome measurement in psychiatric research.
Area of Science:
- Artificial Intelligence in Healthcare
- Natural Language Processing for Clinical Data
- Psychiatric Outcome Measurement
Background:
- Real-world psychiatric care exhibits significant variability in patient presentation and treatment outcomes.
- Systematic outcome measurement is crucial for effective psychiatric care and research.
- The Clinical Global Impression-Severity (CGI-S) scale is a standard clinician-rated measure, but its use in routine care is limited.
Purpose of the Study:
- To assess the feasibility of using Large Language Models (LLMs) to extract CGI-S scores from electronic health record (EHR) clinical notes.
- To compare the performance of different LLM architectures and prompting strategies in estimating CGI-S scores for major depressive disorder (MDD) patients.
- To evaluate the agreement between LLM-generated CGI-S scores and expert clinician ratings.
Main Methods:
- Extracted psychiatrist-authored clinical notes for 77 MDD patients from the Johns Hopkins EHR.
- Three psychiatrists independently rated the CGI-S scores using a validated depression-specific rubric, establishing high inter-rater reliability (κ=0.77-0.78).
- Evaluated GPT-4o (zero-shot and few-shot prompting) and Llama-4 (zero-shot prompting) against human ratings, analyzing agreement based on note characteristics and patient demographics.
Main Results:
- GPT-4o with zero-shot prompting achieved the highest agreement with average human ratings (κ=0.85), outperforming Llama-4 (κ=0.70).
- Few-shot prompting did not enhance GPT-4o's performance.
- Model agreement was not significantly affected by patient demographics or note length percentage, but was lower for shorter notes (κ=0.72 vs. 0.92 for longer notes).
Conclusions:
- LLMs can reliably estimate clinician-rated CGI-S scores from psychiatric notes for MDD patients, achieving agreement levels comparable to expert inter-rater reliability.
- GPT-4o demonstrated superior performance over an open-source alternative, highlighting the impact of model architecture.
- This automated approach holds promise for scalable outcome measurement in research and could facilitate measurement-based care implementation in clinical practice.