Related Experiment Video
Updated: Jan 7, 2026

Rup (RNA-seq Usability Assessment Pipeline) - Quality Control for Bulk RNA-seq Experiments in Eukaryotes
Published on: November 7, 2025
ChatGPT and reference intervals: a comparative analysis of repeatability in GPT-3.5 Turbo, GPT-4, and GPT-4o
Annika Meyer1,2, Edgar Schömig3, Thomas Streichert2
1Department of Anesthesiology and Operative Intensive Care, Faculty of Medicine and University Hospital, University Hospital Cologne, Cologne, Germany.
Background:
Large language models such as ChatGPT hold promise as rapid "curbside consultation" tools in laboratory medicine. However, their ability to generate consistent and clinically reliable reference intervals-particularly in the absence of contextual clinical information-remains uncertain.
Method:
This cross-sectional study evaluated whether three versions of ChatGPT (GPT-3.5-Turbo, GPT-4, GPT-4o) maintain repeatable reference-interval outputs when the prompt intentionally omits the interval, using reference interval variability as a stress-test for model consistency. Standardized prompts were submitted through 726,000 chatbot requests. A total of 246,842 reference intervals across 47 laboratory parameters were then analyzed for consistency using the coefficient of variation (CV) and regression models.
Results:
On average, the chatbots exhibited a CV of 26.50% (IQR: 7.35-129.01%) for the lower limit and 15.82% (IQR: 4.50-45.30%) for the upper limit upon repetition. GPT-4 and GPT-4o demonstrated significantly lower CVs compared to GPT-3.5-Turbo. Reference intervals for poorly standardized parameters were particularly inconsistent across lower (β: 0.6; 95% CI: 0.35 to 0.86; p < 0.001) and upper limit (β: 0.5; 95% CI: 0.28 to 0.71; p < 0.001), while unit expressions also showed variability.
Conclusion:
While the newer ChatGPT versions tested demonstrate improved repeatability, diagnostically unacceptable variability persists, particularly for poorly standardized analytes. Mitigating this requires thoughtful prompt design (e.g., mandatory inclusion of reference intervals), global harmonization of laboratory standards, further model refinement, and robust regulatory oversight. Until then, AI chatbots should be restricted to professional use and trained to refuse laboratory interpretation when reference intervals are not provided by the user.
More Related Videos
08:30Intraperitoneal Glucose Tolerance Test, Measurement of Lung Function, and Fixation of the Lung to Study the Impact of Obesity and Impaired Metabolism on Pulmonary Outcomes
Published on: March 15, 2018
09:30Pre-Implantation Genetic Testing for Aneuploidy on a Semiconductor Based Next-Generation Sequencing Platform
Published on: August 17, 2022
Related Concept Videos
Bioequivalence Data: Statistical Interpretation
Quantifying and Rejecting Outliers: The Grubbs Test
Comparing Experimental Results: Student's t-Test
Improving Translational Accuracy
Improving Translational Accuracy
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...