Related Experiment Video
Updated: Jan 7, 2026

Author Spotlight: A Novel Protocol for Intracameral Injections to Enhance Precision in Rodent Ophthalmology
Published on: May 31, 2024
Comparison of ChatGPT-4o and DeepSeek R1 in the Management of Ophthalmological Emergencies-An Analysis of Ten
Dominik Knebel1, Siegfried Priglinger1, Benedikt Schworm1
1Department of Ophthalmology, University Hospital, LMU Munich, Mathildenstraße 8, 80336 Munich, Germany.
Abstract:
Background: Generative artificial intelligence (AI) applications have gained increasing popularity in recent years and are used by an ever-increasing number of people on a day-to-day basis. While the performance of the earlier-generation generative AI ChatGPT-3.5 in the context of ophthalmologic emergencies has been previously assessed, the purpose of this study is to analyze the performance of the newer-generation generative AIs DeepSeek R1 (Hangzhou DeepSeek Artificial Intelligence Co., Ltd., Hangzhou, China) and ChatGPT-4o (OpenAI Inc., San Francisco, CA, USA) in the context of diagnosis, triage and prehospital management of ophthalmological emergencies. Methods: Ten previously published fictional case vignettes representing queries in the English language of patients experiencing acute ophthalmological symptoms were entered into the generative AIs DeepSeek R1 and ChatGPT-4o. The interaction with the generative AIs followed a previously described structured interaction path. In a random order, each case vignette was entered into separate chats five times, producing a total of 50 answers from each generative AI. Each answer was analyzed according to a previously published manual. Results: We observed better values for DeepSeek R1 compared to ChatGPT-4o in terms of treatment accuracy (60% compared to 50%), the share of answers containing wrong (46% compared to 60%) or conflicting information (30% compared to 40%), the share of answers that correctly captured the overall severity of symptoms (98% compared to 78%), as well as the share of potentially harmful answers (38% compared to 50%). Moreover, DeepSeek R1 more frequently provided a single diagnosis (20% compared to 16%) and specific treatment advice (42% compared to 20%) than ChatGPT-4o. Both generative AIs showed a diagnostic accuracy of 100%, i.e., whenever they provided a single diagnosis, this was indeed the correct diagnosis. In terms of triage accuracy, ChatGPT-4o performed slightly better than DeepSeek R1 (73% compared to 66%). In contrast to DeepSeek R1, which never directed questions back at the user, ChatGPT-4o always did. The direction of questions at the user enables dialogues with ChatGPT-4o that more closely resemble the actual taking of a patient's history. However, DeepSeek R1 seems to perform better compared to ChatGPT-4o in terms of several important content-related metrics and has been shown to be more cost-effective in other studies. Conclusions: Both newer-generation generative AIs constitute remarkable milestones in the development of generative artificial intelligence. However, since potentially harmful recommendations were observed with both models, we currently do not recommend their use as sole source of information on ophthalmological emergencies for laypersons.
Related Concept Videos
Glaucoma: Overview
Angle Closure Glaucoma: Treatment
Open Angle Glaucoma: Treatment
Drugs such as carbonic anhydrase inhibitors, α2- and...

