Related Experiment Video
Updated: Jan 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of large language models on veterinary undergraduate multiple-choice examinations: a comparative
Santiago Alonso Sousa1, Syed Saad Ul Hassan Bukhari1, Paulo Vinicius Steagall1,2
1Department of Veterinary Clinical Sciences, Jockey Club College of Veterinary Medicine and Life Sciences, City University of Hong Kong, Kowloon, Hong Kong SAR, China.
Advanced artificial intelligence, specifically large language models (LLMs), show strong potential in veterinary education. ChatGPT models demonstrated superior performance on veterinary exams, highlighting LLMs as valuable assessment tools.
Area of Science:
- Veterinary Medicine
- Artificial Intelligence
- Educational Technology
Background:
- The application of artificial intelligence (AI), particularly large language models (LLMs), in veterinary education and practice is emerging.
- However, their efficacy in specialized veterinary contexts requires further investigation.
Purpose of the Study:
- To conduct a comparative performance evaluation of nine advanced LLMs on veterinary multiple-choice questions (MCQs).
- To identify factors influencing LLM performance in veterinary assessments.
Main Methods:
- Nine LLMs were tested on 250 MCQs from a veterinary undergraduate final examination.
- Questions covered diverse species, clinical topics, reasoning stages, and included text- and image-based formats.
Main Results:
- ChatGPT models (o1Pro and 4.5) achieved the highest accuracy (90.4% and 90.8%), while Kimi 1.5 performed lowest (64.8%).
- Performance decreased with question difficulty and was lower for image-based questions.
- OpenAI models showed enhanced visual interpretation capabilities.
Conclusions:
- LLMs show significant promise as supportive tools for quality assurance in veterinary assessment design.
- Question difficulty, format, and domain-specific training data are key factors affecting LLM performance.
Related Concept Videos
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Improving Translational Accuracy
Improving Translational Accuracy
