Related Experiment Video
Updated: Jan 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated Safety Plan Scoring in Outpatient Mental Health Settings Using Large Language Models: Exploratory Study
Hayoung K Donnelly1,2, Gregory K Brown1, Kelly L Green1
1Department of Psychiatry, University of Pennsylvania, Philadelphia, PA, United States.
Background:
The safety planning intervention (SPI) is a suicide prevention intervention that results in a written plan to help patients reduce suicide risk. High-quality safety plans-that is, those that are the most complete, personalized, and specific-are more effective in reducing suicide risk. Measuring SPI quality is labor-intensive, which means that clinicians rarely get specific, actionable feedback on their use of the SPI.
Objective:
This study aimed to develop the Safety Plan Fidelity Rater, an automated tool that assesses the quality of written safety plans leveraging 3 large language models (LLMs)-GPT-4, LLaMA 3, and o3-mini.
Methods:
Using 266 deidentified safety plans from outpatient mental health settings in New York, LLMs analyzed four key steps: warning signs, internal coping strategies, making the environment safe, and reasons for living. We compared the predictive performance of the three LLMs, optimizing scoring systems, prompts, and parameters.
Results:
Findings showed that LLaMA 3 and o3-mini outperformed GPT-4, with different step-specific scoring systems recommended based on weighted F1-scores.
Conclusions:
These findings highlight LLMs' potential to provide clinicians with timely and accurate feedback on safety plan quality, which could greatly improve its implementation in community practice.
