Related Experiment Video
Updated: Jan 15, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
Evaluating the Efficacy of AI-Based Interactive Assessments Using Large Language Models for Depression Screening:
Zheng Jin1, Jiaxing Hu2, Dandan Bi1
1International Joint Laboratory of Behavior and Cognitive Science, Zhengzhou Normal University, Zhengzhou, Henan, China.
JMIR Formative Research
|January 13, 2026
Summary
This study shows an automated depression assessment using large language models is feasible, demonstrating high accuracy and user satisfaction compared to traditional methods.
Area of Science:
- Psychological assessment
- Natural Language Processing (NLP)
- Artificial Intelligence (AI) in healthcare
Background:
- Traditional rating scales for psychological assessment have been used for over a century.
- Large language models (LLMs) offer new potential for transforming psychological evaluation.
- This study addresses the need for novel approaches in assessing depressive symptoms.
Purpose of the Study:
- To develop and validate an automated assessment paradigm for depressive symptoms.
- To integrate NLP with conventional measurement tools for psychological evaluation.
- To explore the feasibility of LLM-based tools in clinical practice.
Main Methods:
- 115 participants completed a custom ChatGPT interface for the Beck Depression Inventory Fast Screen (BDI-FS-GPT) and the Patient Health Questionnaire-9 (PHQ-9).
- Statistical analyses included Spearman correlation, Cohen κ for diagnostic agreement, and area under the curve (AUC) for accuracy.
- Comparison was made between the automated tool, the PHQ-9, and clinical diagnosis.
Main Results:
- The BDI-FS-GPT showed moderate correlation with PHQ-9 scores and substantial agreement with clinical diagnosis (κ=0.72).
- The automated tool achieved excellent diagnostic accuracy (AUC=0.953), outperforming the PHQ-9 (AUC=0.859).
- Participants reported significantly higher satisfaction with the automated assessment.
Conclusions:
- An automated assessment paradigm integrating NLP shows preliminary feasibility for psychological evaluation.
- This approach combines interactivity and personalization with psychometric rigor.
- Further validation in larger, diverse studies is warranted as LLM technology advances.