Related Experiment Video
Updated: Aug 6, 2026

Diagnosis of Neoplasia in Barrett’s Esophagus using Vital-dye Enhanced Fluorescence Imaging
Published on: May 11, 2014
Evaluating ChatGPT-5 for Detection of Barrett's Esophagus and Grading of Esophagitis: A Multiclass Endoscopic Image
Hamza R Khan1, Ibraheem Mirza1, Owais M Aftab2
1Division of Gastroenterology & Advanced Endoscopy Rutgers New Jersey Medical School Newark New Jersey USA.
Aims:
Gastroesophageal reflux disease can progress to reflux esophagitis and Barrett's esophagus (BE), making accurate endoscopic diagnosis important. Artificial intelligence tools like ChatGPT-5 may assist image interpretation, though data on newer large language models remains limited. This study evaluated ChatGPT-5 for BE detection and LA esophagitis severity classification.
Methods And Results:
Endoscopic images from the HyperKvasir dataset were analyzed, including BE, esophagitis A, esophagitis B-D, and normal Z-line images. Four standardized prompts were assessed: (1) BE versus normal, (2) esophagitis versus normal, (3) LA-grade severity (A vs. B-D), and (4) BE versus severe esophagitis. ChatGPT-5 was evaluated in auto mode. Two investigators analyzed 640 unique images, yielding 1280 evaluations. Sensitivity, specificity, positive/negative predictive values (PPV/NPV), F1 scores, and accuracy were calculated. In binary tasks, sensitivity was highest for severe esophagitis (B-D) (0.774). Binary accuracy was similar across tasks (~0.64), with the highest for severe esophagitis (0.655). In three-class analyses, severe esophagitis performed best (sensitivity 0.506, specificity 0.761, accuracy 0.438), with performance improving alongside disease severity. PPV trended higher for severe esophagitis compared with normal mucosa (p ~ 0.03), though significance was not retained after multiple-comparison correction. In the BE-severe esophagitis-normal comparison, severe esophagitis achieved the highest sensitivity (0.590), specificity (0.790), and accuracy (0.521). PPV trended higher for severe esophagitis versus normal mucosa (p ~ 0.05), though significance was not maintained after multiplicity adjustment. NPVs exceeded PPVs across all paradigms.
Conclusion:
ChatGPT-5 demonstrated moderate performance for esophageal image interpretation, performing best for severe esophagitis and worst for mild esophagitis/BE. Binary prompting outperformed multiclass formats, and the model functioned better as a rule-out tool.
Related Concept Videos
Barrett Esophagus-I: Introduction
This constant acid exposure transforms the esophagus's pink mucosal lining (stratified squamous epithelium) into a type of lining more similar...
Barrett Esophagus-II: Clinical Manifestations and Management
To diagnose Barrett's esophagus, healthcare providers often recommend an endoscopy for those showing symptoms of acid reflux. The procedure entails...
Endoscopic Procedures I: Esophagogastroduodenoscopy
During an EGD, the endoscope can be used to:
