Related Experiment Video
Updated: Jul 16, 2026

07:15
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
6.8K
Two artificial intelligence models underperform on examinations in a veterinary curriculum.
Journal of the American Veterinary Medical Association
|February 21, 2024
Summary
Artificial intelligence (AI) models like ChatGPT show limitations in veterinary science knowledge. GPT-4.0 performed better than GPT-3.5 but still scored lower than veterinary students, indicating cautious use is advised.
Area of Science:
- Veterinary Education
- Artificial Intelligence in Education
Background:
- Artificial intelligence (AI) and large language models (LLMs) offer new educational possibilities.
- Understanding AI's knowledge in veterinary science is crucial for educators and students.
- ChatGPT platforms (GPT-3.5 and GPT-4.0) require evaluation for their accuracy and consistency.
Purpose of the Study:
- To assess the knowledge level and response consistency of ChatGPT GPT-3.5 and GPT-4.0.
- To compare AI model performance against veterinary student knowledge.
Main Methods:
- Utilized 495 multiple-choice and true/false questions from third-year veterinary courses.
- Entered questions three times into GPT-3.5 and GPT-4.0, recording answers.
- Compared AI-generated answers with faculty-provided correct answers.
Main Results:
- GPT-3.5 achieved 55% performance; GPT-4.0 achieved 77% performance.
- Both AI platforms scored significantly lower than veterinary students (86%).
- GPT-4.0 demonstrated significantly improved performance over GPT-3.5 (P < .05).
Conclusions:
- Current AI models like ChatGPT exhibit knowledge gaps in veterinary science.
- Veterinary educators and students should exercise caution when using AI platforms for information retrieval.
- Further research is needed to improve AI accuracy in specialized scientific domains.

