Related Experiment Video
Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Integrating expert knowledge into large language models improves performance for psychiatric reasoning and diagnosis
Karthik V Sarma1, Kaitlin E Hanss2, Andrew J M Halls2
1Department of Psychiatry and Behavioral Sciences, University of California San Francisco, 675 18th Street, San Francisco, CA 94143, USA; Bakar Computational Health Sciences Institute, University of California San Francisco, 550 16th Street, San Francisco, CA 94143, USA.
Integrating expert reasoning with large language models (LLMs) improved psychiatric diagnostic accuracy. Decision trees reduced overdiagnosis, enhancing LLM performance in behavioral health settings.
Area of Science:
- Artificial Intelligence in Medicine
- Computational Psychiatry
- Clinical Decision Support Systems
Background:
- Large language models (LLMs) show potential in assisting with psychiatric diagnoses.
- Evaluating LLM performance in clinical settings requires rigorous assessment.
- The impact of structured clinical reasoning on LLM diagnostic accuracy is not well-defined.
Purpose of the Study:
- To assess the diagnostic performance of common LLMs using clinical case vignettes.
- To investigate the effect of integrating expert-derived diagnostic decision trees on LLM performance.
- To compare direct LLM prompting versus LLM-assisted reasoning for psychiatric diagnosis.
Main Methods:
- Retrieved clinical case vignettes and diagnoses from DSM-5-TR resources.
- Developed and refined diagnostic decision trees for LLM implementation.
- Prompted three LLMs with vignettes, using both direct and decision-tree-guided approaches.
- Evaluated performance using positive predictive value (PPV), sensitivity, and F1 statistic.
Main Results:
- Direct LLM prompting resulted in high sensitivity but low positive predictive value (PPV), indicating overdiagnosis.
- Utilizing decision trees significantly increased PPV (from 40.4% to 65.3%) with a modest decrease in sensitivity (from 76.7% to 70.9% for the best model).
- Decision trees statistically improved PPV and F1 scores in most experiments, while slightly reducing sensitivity.
Conclusions:
- Direct LLM prompting for psychiatric diagnosis leads to overdiagnosis.
- Integrating expert-derived decision trees enhances LLM diagnostic performance by reducing overdiagnosis.
- LLM-assisted reasoning shows promise for improving clinical decision support tools in behavioral health.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Language and Cognition
Reason and Intuition
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Diagnostic and Statistical Manual of Mental Disorders (DSM)
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Cognitivism
Previously dominated by behaviorism, which prioritized observable behaviors and largely ignored mental processes, psychology transformed in the 1950s. Cognitive psychologists argue that understanding how we think and process...