Related Experiment Video
Updated: Jun 17, 2026

07:51
Full-Field Optical Coherence Microscopy for Histology-Like Analysis of Stromal Features in Corneal Grafts
Published on: October 21, 2022
Early Detection of Keratoconus Using Corneal Tomography Image Analysis by GPT-5.4
Noa Kapelushnik1,2, Daniela Kamar1, Jacob Megreli1,2
1Department of Ophthalmology, Rabin Medical Center, Petach Tikva, Israel.
Cornea
|June 16, 2026
Summary
A large language model (GPT-5.4) showed high specificity but low sensitivity in detecting early keratoconus (KC) from corneal tomography images. It is not currently a reliable tool for KC diagnosis or screening.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Imaging
Background:
- Early detection of keratoconus (KC) is crucial for timely intervention and preventing vision loss.
- Corneal tomography provides detailed structural information for KC diagnosis.
- Large language models (LLMs) are emerging as potential tools for medical image analysis.
Purpose of the Study:
- To assess the diagnostic performance of a general-purpose multimodal LLM (GPT-5.4) in identifying early stages of keratoconus.
- To evaluate the limitations of GPT-5.4 when analyzing corneal tomography images for keratoconus detection.
- To determine the accuracy, sensitivity, and specificity of the LLM in classifying normal, subclinical KC, and forme fruste KC (ffKC) eyes.
Main Methods:
- A retrospective analysis of 66 eyes (15 subclinical KC, 18 ffKC, 33 normal controls) using standardized four-map Pentacam corneal tomography images.
- GPT-5.4 was prompted to analyze images only, without clinical or demographic data.
- Diagnostic performance metrics including sensitivity, specificity, accuracy, and ROC analysis were calculated.
Main Results:
- GPT-5.4 achieved high specificity (100%) but low sensitivity (30.3%) for early KC detection, with an overall accuracy of 65.2%.
- The model exhibited a strong bias towards classifying eyes as normal, leading to a high false-negative rate.
- Performance varied by stage: excellent for normal corneas (100%), moderate for subclinical KC (53.3%), and poor for ffKC (11.1%), with most ffKC cases misclassified as normal.
Conclusions:
- GPT-5.4 demonstrates high specificity but insufficient sensitivity for early keratoconus detection using image-only input.
- The current version of GPT-5.4 is not a reliable tool for diagnosing or screening keratoconus.
- Further research is needed to improve LLM performance for ophthalmic image analysis.
