Related Experiment Video
Updated: Jan 7, 2026

In Vivo Methods to Assess Retinal Ganglion Cell and Optic Nerve Function and Structure in Large Animals
Published on: February 26, 2022
Observer-Performance Comparison of ChatGPT-5 and Gemini 2.5 Pro Versus Veterinarians in Canine and Feline Fundus
Büşra Kibar1, Ömer Tarık Orhun2, Tuğçe Kartal3
1Faculty of Veterinary Medicine, Department of Surgery, Aydın Adnan Menderes University, Aydın, Türkiye.
Objective:
To compare two large language models (ChatGPT-5, Gemini 2.5 Pro) with experienced and novice veterinarians on canine and feline fundus cases, and to assess the relationship between perceived case difficulty and diagnostic performance.
Animals Studied:
Forty-three client-owned cases were sampled from 200 ophthalmology records.
Procedure(S):
Each case included signalment, history, and fundus photographs. Two experienced veterinarians, two novice veterinarians, and two LLMs independently selected findings and provided diagnosis from options. Participants rated difficulty (Very Easy-Hard). Group differences were tested with Kruskal-Wallis and Dunn-Bonferroni procedures; associations with difficulty used Spearman's ρ; paired proportions used Cochran's Q with Holm-adjusted McNemar tests.
Results:
Experts achieved the highest accuracies (findings: 73.3% and 61.6%; diagnosis: 86.0% and 66.3%), significantly outperforming LLMs and novices (all adjusted p < 0.05). LLM finding accuracies were 52.0% (ChatGPT-5) and 49.3% (Gemini 2.5 Pro), both above novices (28.3% and 26.9%). LLM diagnosis accuracies were lower (ChatGPT-5: 37.2%, Gemini 2.5 Pro: 37.2%) but still numerically higher than novices (23.1% and 22.5%). Expert accuracy declined with increasing case difficulty, whereas LLM performance was comparatively stable (ChatGPT-5 range 2.37-3.86; Gemini 2.5 Pro 2.00-2.95). Difficulty correlated negatively with Expert 2 totals (ρ = -0.70, p < 0.0001) but not with LLMs (|ρ| ≤ 0.17, p ≥ 0.28).
Conclusions:
Experienced veterinarians are most accurate in fundus interpretation, but their performance declines with increasing difficulty. LLMs, though less accurate, remain stable across cases and outperform novices, indicating value as training or decision-support tools. Future studies should assess whether expert-LLM collaboration enhances accuracy and efficiency.
More Related Videos
10:28Gene Regulation and Targeted Therapy in Gastric Cancer Peritoneal Metastasis: Radiological Findings from Dual Energy CT and PET/CT
Published on: January 22, 2018
04:18Use of Three-Dimensional Imaging Reconstruction Software as a Training Tool for Cranial Vena Cava Venipuncture in the Ferret
Published on: July 15, 2025