Related Experiment Video
Updated: Jun 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Simulated Reasoning and Self-Verification for Psychiatric Diagnosis in Generalist Large Language Models: Comparative
Karthik V Sarma1,2, Kaitlin E Hanss1, Andrew J M Halls1
1Department of Psychiatry and Behavioral Sciences, University of California, San Francisco, 675 18th Street, Box 3134, San Francisco, CA, 94107, United States, 1 415-476-7527.
Simulated reasoning (SR) and self-verification (SV) significantly improved the positive predictive value (PPV) of large language models (LLMs) for psychiatric diagnosis. These methods enhance diagnostic accuracy without compromising sensitivity, offering promising advancements for behavioral health applications.
Area of Science:
- Artificial Intelligence in Medicine
- Computational Psychiatry
- Natural Language Processing
Background:
- Large language models (LLMs) and large reasoning models (LRMs) show potential in psychiatry but have limitations for accurate psychiatric diagnosis.
- Shortcomings and risks identified in LLM performance complicate their use in clinical settings.
- Simulated reasoning (SR) and self-verification (SV) are emerging techniques to enhance LLM efficacy by guiding output with reasoning tokens.
Purpose of the Study:
- To investigate the impact of SR (using LRMs) and SV (using supplemental prompting) on the psychiatric diagnostic performance of LLMs.
- To evaluate whether these advanced inference approaches improve diagnostic accuracy compared to basic prompting.
Main Methods:
- Extracted 106 DSM-5-TR clinical case vignettes and diagnoses.
- Utilized LLMs and LRMs from OpenAI and Google, employing both a basic direct prompting approach and an SV approach with augmented prompts.
- Evaluated diagnostic performance using sensitivity and positive predictive value (PPV), analyzing results with binomial generalized linear mixed models.
Main Results:
- All models and approaches successfully processed vignettes; sensitivity ranged from 0.732 to 0.817, and PPV ranged from 0.534 to 0.779.
- The best performance was achieved by an LRM using SV (sensitivity: 0.782, PPV: 0.779).
- Statistically significant improvements in PPV were observed with both SR and SV (P=.007 for prompt type, P=.009 for model type), while sensitivity showed no significant fixed effects.
Conclusions:
- Both SR and SV significantly enhance the positive predictive value (PPV) of LLMs for psychiatric diagnosis, without negatively impacting sensitivity.
- The addition of manual SV prompts further boosts PPV, even when SR is employed.
- Future applications of language models in behavioral health can benefit from integrating manually crafted reasoning prompts and automated SR techniques.
Related Concept Videos
Diagnostic and Statistical Manual of Mental Disorders (DSM)
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
