Related Experiment Video
Updated: Jun 23, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Explainable AI for mental health emergency returns: integrating large language models with predictive modeling
Abdulaziz Ahmed1,2, Mohammad Saleem1, Mohammed Alzeen1
1Department of Health Services Administration, School of Health Professions, University of Alabama at Birmingham, Birmingham, AL 35233, United States.
JAMIA Open
|June 22, 2026
Summary
Large Language Models (LLMs) integrated with machine learning (ML) modestly improved predicting emergency department (ED) returns for mental health (MH) patients. This hybrid approach significantly enhanced model interpretability, offering better clinical decision support.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Decision Support Systems
- Mental Health Informatics
Background:
- Emergency department (ED) returns for mental health (MH) conditions represent a significant healthcare burden.
- Traditional machine learning (ML) models show potential for predicting these returns but lack clinical interpretability.
- Enhancing the explainability of predictive models is crucial for clinical adoption.
Purpose of the Study:
- To evaluate the integration of Large Language Models (LLMs) with ML methods for predicting 30-day ED returns in MH patients.
- To assess whether this hybrid approach can improve both predictive performance and model interpretability.
- To explore the utility of LLM-derived features and a novel LLM-SHAP framework for enhanced explanations.
Main Methods:
- Retrospective analysis of 42,464 ED visits for 27,904 unique MH patients (2018-2022).
- LLaMA 3 (8B) used for few-shot classification of chief complaints and social determinants of health (SDoH).
- LLM-extracted features integrated into an XGBoost model; LLM-SHAP framework for enhanced interpretability.
Main Results:
- LLM chief complaint classifier achieved high accuracy (0.882) and recall (0.88).
- SDoH classifiers demonstrated strong performance (F1-scores 0.67-0.96).
- Hybrid model improved predictive performance (AUC from 0.74 to 0.76) and significantly enhanced interpretability via LLM-SHAP.
Conclusions:
- Integrating LLMs with ML modestly improves predictive accuracy for ED returns in MH patients.
- This hybrid approach substantially enhances model interpretability, providing contextualized risk explanations.
- The findings suggest a promising pathway for developing actionable and explainable clinical decision support tools in emergency psychiatry.