Related Experiment Video
Updated: Aug 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
The use of large language models in automated depression detection
Sing Hui Ling1, Wesley Chorney2
1School of Medicine, University of Limerick, County Limerick, Ireland.
Background:
Large language models have been evaluated on many healthcare tasks, including depression screening. However, it is unclear whether estimates of performance are accurate, especially in a setting with realistic clinical constraints.
Methods:
We use the publicly-available Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) to give performance estimates of LLMs that respect patient privacy and could be feasibly deployed in a clinical setting. These models are locally run and under 15 billion parameters.
Results:
Accuracy, sensitivity, and specificity of the models we evaluated ranged from 0.233-0.677, 0.041-0.929, 0.000-0.729, respectively. There are significant differences in the performance we observed versus other studies that evaluate commercial models. We also demonstrate poor agreement amongst different LLMs.
Conclusion:
Current performance estimates of LLMs with respect to depression screening are most likely optimistic. When restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a screening programme.