Related Experiment Video
Updated: Jun 24, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
544
Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication
Philip Chung1, Christine T Fong2, Andrew M Walters2
1Department of Anesthesiology, Perioperative & Pain Medicine, Stanford University, Stanford, California.
JAMA Surgery
|June 5, 2024
Summary
Large language models show promise in predicting surgical outcomes like mortality and admissions, but struggle with duration predictions. Their explanatory capabilities may enhance clinical workflows.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Surgical Risk Prediction
Background:
- General-domain large language models (LLMs) offer potential for analyzing electronic health records (EHRs) for clinical decision support.
- Accurate preoperative risk stratification is crucial for optimizing patient care and resource allocation.
Purpose of the Study:
- To evaluate the predictive performance of the GPT-4 Turbo LLM on eight distinct perioperative outcome measures.
- To assess the LLM's ability to predict both categorical and numerical outcomes using patient EHR data.
Main Methods:
- A prognostic study utilizing retrospective EHR data from a quaternary care center.
- GPT-4 Turbo was prompted with patient case and note data to predict outcomes including American Society of Anesthesiologists Physical Status (ASA-PS), admissions, mortality, and length of stay.
- Prompting strategies including original notes, summaries, few-shot, and chain-of-thought were compared.
Main Results:
- The LLM achieved strong predictive performance for categorical outcomes, with F1 scores ranging from 0.50 (ASA-PS) to 0.86 (hospital mortality).
- Predictive accuracy for numerical duration outcomes (PACU, hospital, and ICU length of stay) was poor across all prompting strategies.
- The model demonstrated varying performance based on prompting techniques, with specific strategies yielding better results for certain tasks.
Conclusions:
- General-domain LLMs can assist in perioperative risk stratification for classification tasks but are currently inadequate for predicting numerical duration outcomes.
- The LLM's capacity for generating natural language explanations enhances its potential utility in clinical workflows.
- LLMs may serve as complementary tools to existing risk prediction models, improving clinical decision-making.

