Related Experiment Video
Updated: Jul 9, 2026

Development of a Virtual Reality Assessment of Everyday Living Skills
Published on: April 23, 2014
Benchmarking large language models against practicing clinicians on psychopathological assessment.
Esra Lenz1,2, Joonas Naamanka3,4,5, Wolfgang Trabert6
1Hector Institute for Artificial Intelligence in Psychiatry, Central Institute of Mental Health, Medical Faculty Mannheim, Heidelberg University, Mannheim, Germany. esra.lenz@gmail.com.
Large language models (LLMs) show promise in psychiatric assessment, accurately evaluating simulated interviews against clinicians. LLMs demonstrated distinct error patterns and improved disagreement resolution, warranting further validation.
Area of Science:
- Psychiatry
- Artificial Intelligence
- Computational Linguistics
Background:
- Psychiatric assessment heavily relies on language, making Large Language Models (LLMs) a potential tool.
- Structured, item-level assessments from clinical interviews using LLMs are under-researched.
Purpose of the Study:
- To evaluate the accuracy of LLMs in psychopathological assessment using structured interview data.
- To compare LLM performance against early-career clinicians and explore their error profiles and utility in disagreement resolution.
Main Methods:
- Ten LLMs assessed transcripts of three simulated psychiatric interviews against the 100-item Association for Methodology and Documentation in Psychiatry (AMDP) system.
- LLM performance was benchmarked against 108 early-career clinicians using audiovisual recordings, with expert consensus as the reference.
- Error profiles and simulated disagreement resolutions between LLMs and clinicians were analyzed.
Main Results:
- GPT-5.1 and Gemini-3-Pro-Preview achieved the highest accuracy (0.72), comparable to the 64th percentile of clinicians.
- GPT-5.1 demonstrated strong performance across depression (0.81), mania (0.76), and schizophrenia (0.60) scenarios.
- LLMs flagged more items as 'not assessable' (19.4%) compared to clinicians (11.4%), especially for observation-dependent items.
Conclusions:
- LLMs show potential as a tool for structured psychopathological assessment, achieving accuracy comparable to clinicians in simulated settings.
- LLMs exhibit different error patterns than clinicians, potentially offering complementary insights.
- LLM assistance in disagreement resolution improved accuracy, suggesting their utility in clinical decision support, though validation in real-world settings is crucial.
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
Treatment Strategies for Psychological Disorders
Psychological therapies focus on modifying emotions, thoughts, and behaviors through talking, interpreting, listening, rewarding, challenging, and modeling. Clinical psychologists, counselors, and social workers commonly practice psychotherapy. Clinical...
Self-Report Tests of Personality
Theoretical Approaches to Psychological Disorder
Biological approach
The biological approach posits that internal, organic factors are the primary causes of such disorders. This perspective emphasizes brain structure and function, genetic predispositions, and neurotransmitter imbalances. For example, schizophrenia has been associated with both genetic...
Diagnostic and Statistical Manual of Mental Disorders (DSM)