GPT-4 Underperforms Experts in Detecting IV Fluid Contamination
Nicholas C Spies1, Zita Hubler1, Stephen M Roper1,2
1Department of Pathology and Immunology, Washington University in St. Louis School of Medicine, St. Louis, MO, United States.
Background:
Specimens contaminated with intravenous (IV) fluids are common in clinical laboratories. Current methods for detecting contamination rely on insensitive and workflow-disrupting delta checks or manual technologist review. Herein, we assessed the utility of large language models for detecting contamination by IV crystalloids and compared its performance to multiple, but variably trained healthcare personnel (HCP).
Methods:
Contamination of basic metabolic panels was simulated using 0.9% normal saline (NS), with (n = 30) and without (n = 30) 5% dextrose (D5NS), at mixture ratios of 0.10 and 0.25. A multimodal language model (GPT-4) and a diverse panel of 8 HCP were asked to adjudicate between real and contaminated results. Classification performance, mixture quantification, and confidence was compared by Wilcoxon rank sum.
Results:
The 95% CIs for accuracy were 0.57-0.71 vs 0.73-0.80 for GPT-4 and HCP, respectively, on the NS set and 0.57-0.57 vs 0.73-0.80 on the D5NS set. HCP overestimated severity of contamination in the 0.10 mixture group (95% CI of estimate error, 0.05-0.20) for both fluids, while GPT-4 markedly overestimated the D5NS mixture at both ratios (0.16-0.33 for NS, 0.11-0.35 for D5NS). There was no correlation between reported confidence and likelihood of a correct classification.
Conclusions:
GPT-4 is less accurate than trained HCP for detecting IV fluid contamination of basic metabolic panel results. However, trained individuals were imperfect at identifying contaminated specimens implying the need for novel, automated tools for its detection.
Insights
Large language models like GPT-4 are less accurate than healthcare personnel in detecting intravenous (IV) fluid contamination in lab specimens. Automated tools are still needed for reliable IV fluid contamination detection.
Area of Science:
- Clinical Laboratory Science
- Artificial Intelligence in Healthcare
- Diagnostic Accuracy
Background:
- Intravenous (IV) fluid contamination in clinical laboratory specimens is a frequent issue.
- Current detection methods, such as delta checks and manual review, are often insensitive and disrupt laboratory workflows.
- There is a need for more effective and automated methods to identify IV fluid contamination.
Purpose of the Study:
- To evaluate the effectiveness of large language models (LLMs) in detecting IV crystalloid contamination in laboratory specimens.
- To compare the performance of an LLM (GPT-4) against trained healthcare personnel (HCP) in identifying contaminated samples.
Main Methods:
- Simulated contamination of basic metabolic panels using normal saline (NS) and 5% dextrose in normal saline (D5NS) at varying mixture ratios.
- A multimodal LLM (GPT-4) and a panel of 8 HCP were tasked with distinguishing between genuine and contaminated results.
- Performance metrics including classification accuracy, mixture quantification, and confidence levels were compared using statistical analysis.
Main Results:
- Healthcare personnel demonstrated higher accuracy (95% CIs: 0.73-0.80) than GPT-4 (95% CIs: 0.57-0.71 for NS, 0.57-0.57 for D5NS) in detecting IV fluid contamination.
- HCP tended to overestimate contamination severity in lower mixture ratios, while GPT-4 significantly overestimated contamination with D5NS mixtures.
- No correlation was found between the confidence level reported by either GPT-4 or HCP and the accuracy of their classifications.
Conclusions:
- Large language models, specifically GPT-4, are currently less accurate than trained healthcare personnel for detecting IV fluid contamination in basic metabolic panels.
- The imperfect performance of even trained individuals highlights the ongoing need for novel, automated solutions for reliable contamination detection in clinical laboratories.


