Related Experiment Video
Updated: May 28, 2026

06:16
LipidUNet-Machine Learning-Based Method of Characterization and Quantification of Lipid Deposits Using iPSC-Derived Retinal Pigment Epithelium
Published on: July 28, 2023
Comparative Analysis of General-Purpose vs. Domain-Specific Multimodal Models for Diabetic Retinopathy Classification
Mohammad Iqbal Nouyed1, Mohammad Al-Mamun2, Donald A Adjeroh3
1Department of Microbiology, Immunology and Cell Biology, West Virginia University, Morgantown, WV 26506, USA.
Diagnostics (Basel, Switzerland)
|May 27, 2026
Summary
General-purpose AI models like Gemini 3 and GPT-5.2 show promise in diagnosing diabetic retinopathy, achieving accuracy comparable to specialized models. However, domain-specific models like MedSigLIP offer superior performance for retinal image analysis.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Ophthalmology AI
- Diabetic Retinopathy Detection
Background:
- Multimodal foundation models, both general-purpose and domain-specific, are emerging as powerful tools for medical image analysis.
- Evaluating the diagnostic capabilities of various AI models for diabetic retinopathy classification is crucial for clinical adoption.
Purpose of the Study:
- To assess the classification accuracy of diabetic retinopathy versus normal fundus images using diverse AI models.
- To compare the performance of general-purpose conversational models against specialized ophthalmology models.
Main Methods:
- Zero-shot, few-shot prompting, linear probing, and fine-tuning techniques were applied to evaluate model performance.
- Models tested included general-purpose (Gemini 3 Flash, GPT-5.2, Pixtral-Large), medical-specific (MedGemma-1.5, MedSigLIP), and ophthalmology-specific (RETFound, EyeCLIP).
Main Results:
- MedSigLIP achieved the highest accuracy (94.8%), followed by MedGemma-1.5 (88.2%) and Gemini 3 (88.5%).
- General-purpose models like GPT-5.2 and Gemini 3 demonstrated competitive zero-shot accuracy, comparable to fine-tuned specialized models.
- Domain-specific models generally offered higher accuracy and stability, while general-purpose models provided flexibility and interactive reasoning.
Conclusions:
- A trade-off exists between the specialization of domain-specific models and the flexibility of general-purpose multimodal models.
- General-purpose models offer accessibility and adaptability, serving as valuable complementary tools for retinal disease screening and clinical decision support.
- Specialized models like MedSigLIP currently provide superior accuracy for diabetic retinopathy classification.