Related Experiment Video
Updated: Jan 7, 2026

Rup (RNA-seq Usability Assessment Pipeline) - Quality Control for Bulk RNA-seq Experiments in Eukaryotes
Published on: November 7, 2025
ChatGPT and reference intervals: a comparative analysis of repeatability in GPT-3.5 Turbo, GPT-4, and GPT-4o
Annika Meyer1,2, Edgar Schömig3, Thomas Streichert2
1Department of Anesthesiology and Operative Intensive Care, Faculty of Medicine and University Hospital, University Hospital Cologne, Cologne, Germany.
Large language models like ChatGPT show promise in lab medicine but struggle with consistent reference intervals. Newer versions improve, yet variability persists, especially for unstandardized tests.
Area of Science:
- Artificial Intelligence in Healthcare
- Laboratory Medicine and Diagnostics
- Clinical Pathology and Informatics
Background:
- Large language models (LLMs) offer potential for rapid clinical consultation in laboratory medicine.
- Uncertainty exists regarding the consistency and clinical reliability of reference intervals generated by LLMs, especially without clinical context.
Purpose of the Study:
- To evaluate the repeatability of reference interval outputs from three ChatGPT versions (GPT-3.5-Turbo, GPT-4, GPT-4o).
- To assess model consistency by using reference interval variability as a stress test when prompts omit interval information.
Main Methods:
- A cross-sectional study involving 726,000 chatbot requests with standardized prompts.
- Analysis of 246,842 reference intervals across 47 laboratory parameters for consistency.
- Statistical analysis using coefficient of variation (CV) and regression models to assess variability.
Main Results:
- Average CVs for reference intervals were 26.50% (lower limit) and 15.82% (upper limit).
- GPT-4 and GPT-4o demonstrated significantly lower CVs than GPT-3.5-Turbo.
- Inconsistent outputs were notable for poorly standardized parameters and varied unit expressions.
Conclusions:
- While newer ChatGPT versions show improved repeatability, diagnostically unacceptable variability remains, particularly for unstandardized analytes.
- Thoughtful prompt design, global standardization of lab practices, model refinement, and regulatory oversight are crucial.
- Current AI chatbots should be limited to professional use and trained to decline interpretation without provided reference intervals.
More Related Videos
08:30Intraperitoneal Glucose Tolerance Test, Measurement of Lung Function, and Fixation of the Lung to Study the Impact of Obesity and Impaired Metabolism on Pulmonary Outcomes
Published on: March 15, 2018
09:30Pre-Implantation Genetic Testing for Aneuploidy on a Semiconductor Based Next-Generation Sequencing Platform
Published on: August 17, 2022
Related Concept Videos
Bioequivalence Data: Statistical Interpretation
Quantifying and Rejecting Outliers: The Grubbs Test
Comparing Experimental Results: Student's t-Test
Improving Translational Accuracy
Improving Translational Accuracy
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...