進化するコンサルテーション:大規模言語モデルを用いた眼科診断パフォーマンスの向上
Taiga Inooka1, Hikaru Ota1, Yosuke Taki1
1Department of Ophthalmology, Nagoya University Graduate School of Medicine, Nagoya, Japan.
Objective:
Artificial intelligence-powered large language models (LLMs) are increasingly applied in health care. However, studies in ophthalmology assessing whether LLMs can improve the accuracy of complex differential diagnoses in clinical cases, or which levels of clinical experience benefit most from their use, remain lacking. This study assessed the effectiveness of ChatGPT-4o, an LLM-driven chatbot, in enhancing ophthalmologists' clinical reasoning using original scenarios.
Design:
Prospective study.
Subjects:
Ten original ophthalmic clinical scenarios with open-ended questions were developed, covering the following subspecialties: oculoplastic and orbital disease, glaucoma, inherited retinal disease, macular disease, neuro-ophthalmology, ocular surface, pediatric ophthalmology, retinal vascular disease, strabismus, and uveitis.
Methods:
Responses to each clinical scenario were collected from 20 ophthalmologists (10 residents and 10 board-certified ophthalmologists) and ChatGPT-4o. Ophthalmologists subsequently revised their answers with assistance from ChatGPT-4o. All responses were anonymized and independently evaluated by 3 attending ophthalmologists based on 4 metrics: coherency, factuality, comprehensiveness, and safety (each on a 5-point scale).
Main Outcome Measures:
The median total scores for each group in coherency, factuality, comprehensiveness, and safety (maximum of 15 points each).
Results:
Assistance from ChatGPT-4o significantly improved evaluation scores for coherency, comprehensiveness, and safety among both residents and board-certified ophthalmologists (all, P < 0.001). However, factuality scores showed no significant improvements (P = 0.114 and 0.839, respectively). Although ChatGPT-4o assistance increased citation frequency (residents: 0.24-0.98 per response, board-certified ophthalmologists: 0.12-0.68 per response, both P < 0.05), approximately 44% of these additional citations were identified as hallucinated references, nonexistent, or incorrect citations. Notably, ChatGPT-4o assistance led to a significant increase in variability for factuality and safety scores in both groups (Brown-Forsythe test, all P < 0.05), whereas it decreased variability for coherency and comprehensiveness, with the reduction statistically significant among residents (P = 0.008 and P = 0.006, respectively).
Conclusions:
ChatGPT-4o effectively enhanced diagnostic reasoning and response quality, particularly among ophthalmology residents. However, successful integration into clinical education and practice requires careful management of increased variability in factuality and safety. This issue could be addressed by implementing strategies such as advanced retrieval-augmented generation systems to ensure the provision of accurate and safe clinical information.
Financial Disclosures:
Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
さらに関連する動画
04:48Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
関連する概念動画
Improving Translational Accuracy
Improving Translational Accuracy
