Related Experiment Video
Updated: Sep 8, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
680
Assessment of Differential Diagnoses for Oculoplastics Cases Produced by Large Language Models
Jeffrey C Peterson1, Sruti S Rachapudi1, Sasha Hubschman1
1Department of Ophthalmology and Visual Sciences, University of Illinois at Chicago, Chicago, Illinois.
Ophthalmic Plastic and Reconstructive Surgery
|August 11, 2025
Summary
Large language models (LLMs) show promise for oculoplastic diagnosis. OcuSmart/EyeGPT and Claude 3.5 excelled in accuracy and recall, while Gemini offered precision in differential diagnoses.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Diagnostics
Background:
- Accurate differential diagnoses are crucial for effective oculoplastic patient care.
- Evaluating the diagnostic capabilities of emerging AI technologies like large language models (LLMs) is essential.
Purpose of the Study:
- To assess the accuracy of six distinct large language models in generating differential diagnoses for oculoplastic conditions.
- To compare the performance of various LLMs against expert-curated diagnoses.
Main Methods:
- Twenty oculoplastic cases from EyeRounds.org were used to generate differential diagnoses with six LLMs: ChatGPT 3.5, ChatGPT 4.0, OcuSmart/EyeGPT, Google Gemini 1.5, Claude 3.5, and Microsoft CoPilot.
- LLM outputs were evaluated based on top diagnosis match rate, inclusion of correct diagnosis, recall, and precision compared to expert differentials.
Main Results:
- OcuSmart/EyeGPT achieved the highest top diagnosis match rate (85%).
- Claude 3.5 showed the highest correct diagnosis inclusion and recall (100% and 55%, respectively).
- Google Gemini provided the most precise differentials (43%), though Claude 3.5 generated less concise lists. Performance varied by case type.
Conclusions:
- LLMs demonstrate significant potential for aiding in oculoplastic case diagnosis.
- OcuSmart/EyeGPT and Claude 3.5 are strong performers for diagnosis and recall, while ChatGPT 3.5, OcuSmart/EyeGPT, and Gemini offer concise differentials.
- Further validation and integration studies are needed to incorporate LLMs into clinical practice.
Related Concept Videos
Prosopagnosia
249
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...
249
Assessment of Airway, Skin Color, and Use of Accessory Muscles
1.1K
A thorough assessment of respiratory health is paramount in clinical settings to identify and manage respiratory distress and ensure adequate oxygenation. This article elaborates on the critical aspects of respiratory evaluation, including airway assessment, skin color examination, and the observation of accessory muscle use, which are integral to effectively diagnosing and managing patients with respiratory conditions.
Introduction
The initial evaluation of a patient's respiratory system...
Introduction
The initial evaluation of a patient's respiratory system...
1.1K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K

