Related Experiment Video
Updated: Jun 14, 2025

03:19
Author Spotlight: Developing a Bedside Protocol for Kidney and Genitourinary Ultrasonography
Published on: June 21, 2024
1.0K
Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones
Medrxiv : the Preprint Server for Health Sciences
|August 30, 2024
Summary
Large language models like GPT-4 can accurately analyze emergency department reports to identify kidney stones, outperforming traditional methods. GPT-4 showed no racial or gender bias, demonstrating its potential for clinical applications.
Area of Science:
- Artificial Intelligence
- Clinical Informatics
- Medical Natural Language Processing
Background:
- Large language models (LLMs) show promise for clinical applications.
- Their ability to analyze emergency department (ED) reports for specific conditions like kidney stones is largely unexplored.
- This study evaluates LLMs against traditional machine learning and ICD codes.
Purpose of the Study:
- To investigate the performance of multiple LLMs (GPT-4, GPT-3.5, Llama-2) in analyzing ED reports for symptomatic kidney stones.
- To compare LLM performance with traditional machine learning models and ICD code-based systems.
- To assess fairness and bias in LLM predictions concerning race and gender.
Main Methods:
- Utilized a manually annotated dataset of ED reports.
- Employed prompt optimization, zero/few-shot prompting, fine-tuning, and prompt augmentation for LLMs.
- Performed fairness assessments and bias mitigation.
- Clinical expert evaluated GPT-4's explanation soundness and factual correctness.
- Compared LLMs against logistic regression, XGBoost, LGBM, and an ICD code baseline.
Main Results:
- GPT-4 (macro-F1=0.833) and GPT-3.5 (macro-F1=0.796) significantly outperformed the ICD baseline (macro-F1=0.71).
- Fine-tuning improved GPT-3.5 performance.
- Including demographic and prior disease history enhanced LLM accuracy.
- GPT-4 demonstrated no racial or gender bias, unlike GPT-3.5.
- GPT-4 explanations indicated strong clinical text understanding and reasoning.
Conclusions:
- LLMs, particularly GPT-4, show significant potential for accurately analyzing clinical notes in emergency departments.
- GPT-4 demonstrates robust performance, fairness, and sound clinical reasoning capabilities.
- Further development and validation are warranted for integrating LLMs into clinical workflows for conditions like kidney stones.

