Related Experiment Video
Updated: Jun 30, 2026

07:31
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Comparative Evaluation of Pretrained Large Language Models for Suicide Risk Prediction from Clinical Notes in U.S.
Joshua Levy1, Maxwell Levis2, Monica Dimambro2
1Department of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA, USA.
Medrxiv : the Preprint Server for Health Sciences
|June 29, 2026
Summary
Large language models (LLMs) show promise in predicting veteran suicide risk by analyzing clinical notes. These advanced NLP methods outperform traditional techniques, potentially improving early identification and intervention for at-risk individuals.
Area of Science:
- Computational psychiatry
- Natural Language Processing (NLP) in healthcare
- Veterans' mental health research
Background:
- Suicide is a critical concern among U.S. veterans, necessitating improved predictive models.
- Electronic Health Records (EHRs) and clinical narratives offer data for risk assessment.
- Large Language Models (LLMs) present novel methods for analyzing unstructured clinical text.
Purpose of the Study:
- To compare the effectiveness of LLM-based text representations against traditional bag-of-words (BoW) models for predicting veteran suicide risk.
- To evaluate model performance across different risk tiers and time windows using clinical notes.
Main Methods:
- Utilized clinical notes from 27,241 veterans within the Veterans Health Administration.
- Compared pretrained LLMs with BoW representations for suicide risk prediction.
- Stratified patients by the Recovery Engagement and Coordination for Health-Veterans Enhanced Treatment (REACH-VET) risk tiers and analyzed various look-back periods.
Main Results:
- LLM representations surpassed BoW in seven of nine evaluated combinations, with a maximum AUROC of 0.644 using text alone.
- Integrating structured clinical data significantly enhanced predictive performance, reaching an AUROC of 0.748.
- Analysis highlighted the importance of suicide-related language in notes, particularly within 30 days of the outcome for high-risk patients.
Conclusions:
- Pretrained LLMs effectively extract clinically relevant information from narrative documentation for suicide risk prediction.
- LLM-based approaches offer a robust foundation for developing more accurate suicide risk prediction models in clinical settings.
- Future research should focus on adapting LLMs to diverse clinical contexts and temporal dynamics to further refine prediction accuracy.