Related Experiment Video
Updated: May 23, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models are less effective at clinical prediction tasks than locally trained machine learning models
Katherine E Brown1, Chao Yan1, Zhuohang Li2
1Department of Biomedical Informatics, Vanderbilt University Medical Center (VUMC), Nashville, TN 37203, United States.
Traditional machine learning (ML) significantly outperforms large language models (LLMs) like GPT-3.5 and GPT-4 for clinical prediction tasks using electronic health records (EHRs). LLMs show limitations in performance, calibration, and privacy protection compared to traditional ML methods.
Area of Science:
- Artificial Intelligence in Healthcare
- Clinical Predictive Modeling
- Electronic Health Records (EHR) Data Analysis
Background:
- Large Language Models (LLMs) are increasingly explored for various applications, including healthcare.
- Traditional Machine Learning (ML) methods are established for clinical prediction using Electronic Health Records (EHRs).
- Evaluating LLMs as potential substitutes for traditional ML in clinical settings is crucial for advancing healthcare technology.
Purpose of the Study:
- To assess the efficacy of current LLMs (GPT-3.5, GPT-4) as clinical predictors compared to traditional ML.
- To investigate factors influencing LLM adoption in clinical prediction: performance, calibration, fairness, and privacy resilience.
- To analyze the impact of data generalization for privacy on model performance.
Main Methods:
- Comparative analysis of GPT-3.5, GPT-4, and gradient-boosting trees (traditional ML) on EHR data from VUMC and MIMIC IV.
- Performance measured by Area Under the Receiver Operating Characteristic (AUROC) and Brier Score for calibration.
- Fairness evaluated using equalized odds and statistical parity across demographic groups.
- Impact of in-context learning and data generalization for privacy on AUROC was assessed.
Main Results:
- Traditional ML demonstrated substantially superior predictive performance (AUROC: 0.847, 0.894) compared to GPT-3.5 (AUROC: 0.537, 0.517) and GPT-4 (AUROC: 0.629, 0.602).
- Traditional ML also exhibited significantly better output probability calibration (Brier Score: 0.134, 0.042) than GPT-3.5 (0.384, 0.06) and GPT-4 (0.251, 0.219).
- GPT-4 showed better fairness across demographic groups but at the expense of reduced model performance.
Conclusions:
- Non-fine-tuned LLMs are currently less effective and robust than locally trained ML for clinical prediction using EHR data.
- Traditional ML models are more robust in generalizing demographic information for privacy protection.
- While current LLMs lag behind traditional ML, ongoing advancements suggest potential for future clinical applications.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Related Concept Videos
Language and Cognition
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Improving Translational Accuracy