Related Experiment Video
Updated: Jul 6, 2025

07:11
Assessing Early Stage Open-Angle Glaucoma in Patients by Isolated-Check Visual Evoked Potential
Published on: May 25, 2020
6.4K
Exploring the Potential of ChatGPT-4 in Predicting Refractive Surgery Categorizations: Comparative Study.
Aleksandar Ćirković1, Toam Katz2
1Care Vision Germany, Ltd, Nuremberg, Germany.
JMIR Formative Research
|December 28, 2023
Summary
ChatGPT-4 shows potential for refractive surgery patient precategorization, achieving notable agreement with clinician assessments. Further research is needed to address model variability and improve accuracy for clinical implementation.
Area of Science:
- Ophthalmology and Artificial Intelligence
- Medical Decision Support Systems
- Refractive Surgery Patient Assessment
Background:
- Artificial intelligence (AI), including machine learning (ML), is advancing refractive surgery by improving patient risk assessment and workflow.
- Large language models (LLMs) like ChatGPT-4 offer potential for diverse applications, including refractive surgery decision-making, but their real-world efficacy is unproven.
Purpose of the Study:
- To explore and validate the capability of ChatGPT-4 in precategorizing refractive surgery patients using standard clinical parameters.
- To compare ChatGPT-4's batch input categorization performance against that of a refractive surgeon.
- To evaluate performance using both binary (suitable/unsuitable for laser surgery) and detailed categorization schemes.
Main Methods:
- Anonymized data from 100 refractive surgery patients, including demographics, refraction, visual acuity, and corneal imaging (Scheimpflug), were analyzed.
- ChatGPT-4's categorizations were compared to those of a refractive surgeon using Cohen κ coefficient, chi-square tests, confusion matrices, accuracy, precision, recall, F1-score, and ROC AUC.
Main Results:
- A statistically significant agreement was observed between ChatGPT-4 and clinician categorizations (Cohen κ = 0.399 for 6 categories, 0.610 for binary).
- Binary categorization yielded higher performance metrics: accuracy 0.88, precision 0.88, recall 0.88, F1-score 0.88, and ROC AUC 0.79.
- The model exhibited temporal instability and response variability, despite a significant association found via chi-square test (χ²₅=94.7, P<.001) for 6 categories.
Conclusions:
- ChatGPT-4 demonstrates potential as a refractive surgery precategorization tool, with promising agreement with expert clinicians.
- Key limitations include reliance on a single rater, small sample size, output instability, and model opacity, necessitating further investigation.
- Future research should standardize prompts and vignettes, identify confounding factors, and compare various LLMs to facilitate validation and clinical integration.
Keywords:
AI-powered algorithmChatGPTChatGPT-4artificial intelligencecategorizationclinicaldata analysisdecision support systemsdecision-makingeHealthhealth informaticslarge language modelmachine learningmedical decision-makingophthalmologypredictive modelingrefractive surgeryrefractive surgical proceduresrisk assessment
