Related Experiment Video
Updated: May 28, 2025

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
2.6K
Integrated visual and text-based analysis of ophthalmology clinical cases using a large language model
Vera Sorin1,2,3, Noa Kapelushnik4,5, Idan Hecht4,6
1Department of Diagnostic Imaging, Chaim Sheba Medical Center, Emek Haela St. 1, 52621, Ramat Gan, Israel. verasrn@gmail.com.
Scientific Reports
|February 10, 2025
Summary
Generative artificial intelligence, specifically multimodal GPT-4V, can analyze ocular images and clinical text for diagnosis. Adding clinical context significantly improved diagnostic accuracy for both AI and physicians.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Diagnostics
Background:
- Ophthalmology diagnosis relies on integrating ocular examination findings with clinical context.
- Generative artificial intelligence (AI) models are increasingly capable of analyzing combined text and visual data.
- Multimodal AI holds potential for enhancing healthcare applications, including medical diagnosis.
Purpose of the Study:
- To evaluate the diagnostic performance of multimodal GPT-4V using integrated ocular images and clinical text.
- To compare the AI's diagnostic accuracy with that of non-ophthalmology physicians.
- To assess the impact of clinical context on diagnostic accuracy for both AI and human evaluators.
Main Methods:
- A retrospective study involving 40 ophthalmology patient cases with ocular images.
- GPT-4V was provided with images alone and then with accompanying clinical text.
- Non-ophthalmology physicians also provided diagnoses with and without clinical context.
- Diagnoses from GPT-4V and physicians were evaluated by board-certified ophthalmologists.
Main Results:
- GPT-4V achieved 47.5% accuracy with images alone and 67.5% with added clinical context.
- Non-ophthalmology physicians achieved 60.0% accuracy without context and 72.5% with context.
- Adding clinical context significantly improved diagnostic accuracy for all participants (p=0.033).
Conclusions:
- Multimodal GPT-4V demonstrates capability in simultaneously analyzing visual and textual data for accurate clinical diagnosis.
- The integration of clinical context enhances diagnostic performance for both AI and human physicians.
- Multimodal large language models show significant promise for advancing ophthalmology patient care and research.

