Related Experiment Video
Updated: Jan 9, 2026

In Vivo Multimodal Imaging and Analysis of Mouse Laser-Induced Choroidal Neovascularization Model
Published on: January 21, 2018
Evaluating the clinical utility of multimodal large language models in rare maculopathy
Melanie D Tran1,2, Evan Walker3, Ines D Nagel1,3
1Retinal Division, Jacobs Retina Center, Shiley Eye Institute, University of California San Diego, 9415 Campus Point Dr, La Jolla, CA, 92093, USA.
Abstract:
This study aimed to assess how multimodal large language models (MLLM) diagnose and differentiate Pentosan Polysulfate (PPS) Maculopathy from other phenotypic mimics. A retrospective review of clinical records and multimodal retinal imaging was conducted with patients from the Shiley Eye Institute and Casey Eye Institute. Four MLLMs (ChatGPT-4o, Claude 3.5 Sonnet, Google Gemini 1.5 Pro, Perplexity Llama 3.1 Sonar/Default) along with human retinal specialists answered prompts based on retinal imaging and demographic data. Performance was evaluated using accuracy, sensitivity and specificity estimates. The study included 126 eyes from 63 patients, with 36 eyes with PPS maculopathy, 50 eyes with Stargardt disease, and 40 eyes with PRPH2-associated multifocal pattern dystrophy. MLLMs showed improved accuracy and sensitivity when answer choices were restricted, with ChatGPT consistently performing best when all imaging modalities were prompted together. The inclusion of demographic data further enhanced performance in prompts with limited answer choices. Human retinal specialist evaluations aligned with MLLM performance trends and also improved with demographic data. While MLLMs show diagnostic potential, further refinement is needed before clinical implementation. These findings highlight the importance of prompt design and demographic data to optimize MLLM performance with retinal imaging modalities.
Insights
Multimodal large language models (MLLMs) show promise in diagnosing Pentosan Polysulfate (PPS) Maculopathy, improving accuracy with specific prompts and demographic data. Further research is needed for clinical use.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Diagnostics
Background:
- Pentosan Polysulfate (PPS) Maculopathy is a condition that can mimic other retinal diseases.
- Accurate differentiation is crucial for appropriate patient management and treatment.
- Multimodal large language models (MLLMs) are emerging tools with potential diagnostic capabilities.
Purpose of the Study:
- To evaluate the diagnostic performance of MLLMs in identifying and differentiating PPS Maculopathy from similar retinal conditions.
- To compare MLLM performance against human retinal specialists.
- To assess the impact of prompt design and demographic data on diagnostic accuracy.
Main Methods:
- Retrospective review of clinical records and multimodal retinal imaging from 63 patients (126 eyes).
- Four MLLMs (ChatGPT-4o, Claude 3.5 Sonnet, Google Gemini 1.5 Pro, Perplexity Llama 3.1 Sonar/Default) and human retinal specialists responded to prompts.
- Performance metrics included accuracy, sensitivity, and specificity, with variations in prompt complexity and data inclusion.
Main Results:
- MLLMs demonstrated improved accuracy and sensitivity when provided with restricted answer choices.
- ChatGPT-4o showed superior performance when all imaging modalities were prompted simultaneously.
- Inclusion of demographic data significantly enhanced MLLM diagnostic performance, particularly with limited choices.
- Human specialist performance trends mirrored MLLMs, also improving with demographic data.
Conclusions:
- MLLMs exhibit diagnostic potential for retinal diseases like PPS Maculopathy.
- Prompt engineering and the integration of demographic data are critical for optimizing MLLM diagnostic accuracy.
- Further refinement and validation are necessary before MLLMs can be clinically implemented for retinal disease diagnosis.

