Related Experiment Video
Updated: Jul 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative diagnostic performance of large language models and clinicians for splenic diseases in dogs and cats
Murat İlgün1, Emre Eren2, Büşra Kibar3
1DVM Graduate, Pethome Veterinary Polyclinic, Istanbul, Turkiye.
Abstract:
Splenic diseases in dogs and cats present significant diagnostic challenges, particularly in differentiating benign from malignant lesions using conventional clinical and imaging modalities. While histopathology remains the gold standard, artificial intelligence, especially large language models (LLMs), offers emerging diagnostic support capabilities. This two-center, retrospective diagnostic-accuracy study compared the performance of two state-of-the-art LLMs (ChatGPT-5 and Gemini 1.5 Pro) with experienced and novice veterinary clinicians in diagnosing splenic diseases in 38 dogs and cats that underwent splenectomy between 2021 and 2025. Each assessor received standardized multimodal case packets, including clinical, laboratory, ultrasonographic, and intraoperative macroscopic data. Histopathology served as the reference standard. Diagnostic performance was evaluated using generalized linear mixed-effects models, malignancy ROC analysis, and information-modality sensitivity assessments. Experts achieved the highest exact specific-entity accuracy (92.1% and 89.5%), followed by ChatGPT-5 (76.3%) and Gemini 1.5 Pro (71.1%), whereas novices performed lowest (57.9% and 52.6%). Upper-category classification was generally higher than exact diagnosis across groups, and adding imaging and macroscopic data improved accuracy for all assessors. Receiver-operating performance followed a consistent gradient (Expert > LLM > Novice), with AUCs of 0.97 for Expert-1, 0.86 for ChatGPT-5, and 0.71 for Novice-2. These findings indicate that modern LLMs, even in zero-shot settings, can provide meaningful diagnostic triage support that exceeds novice performance and, when multimodal inputs are available, approaches but does not demonstrate equivalence to expert-level classification. Prospective multicenter studies and the evaluation of natively multimodal models capable of directly interpreting clinical images may further enhance their clinical applicability in veterinary diagnostic workflows.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025