Related Experiment Video
Updated: May 4, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Accuracy of large language model-based artificial intelligence tools for equine topics
S Aldworth-Yang1, S J Coleman1, K O'Reilly1
1Department of Animal Sciences, College of Agricultural Sciences, Colorado State University, 350 W Pitkin Street, Fort Collins, Colorado 80521, USA.
Background:
Artificial intelligence (AI) platforms are becoming increasingly popular as resources for equine information. However, these platforms generate responses from a wide range of sources and do not always distinguish between fact and opinion.
Aims/Objectives:
The objective of this study was to assess the accuracy and quality of AI-generated answers to equine-related questions. Researchers hypothesized that AI platforms could answer basic equine questions effectively but would perform poorly on complex topics or questions.
Methods:
Forty questions were written covering general horse care, facilities management, nutrition, genetics, and reproduction. Each question was categorized by difficulty level: beginner, intermediate, advanced, or trending. Three AI platforms were tested: ChatGPT (CGPT), Microsoft Copilot (MicCP), and ExtensionBot (ExtBot). Responses were scored for accuracy, relevance, thoroughness, and source quality (5 points each; total 20). Data were analyzed using PROC GLM in SAS (v. 9.4).
Results:
Total score was affected by level (P = 0.002). Intermediate questions had the highest total score (15.95 ± 1.99). Accuracy was affected by platform (P < 0.001), level (P < 0.001), and topic (P = 0.015). CGPT (4.18 ± 0.93) and MicCP (4.08 ± 0.83) outperformed ExtBot (3.26 ± 1.21). Relevance was affected by platform (P = 0.042) and level (P < 0.001). Thoroughness was affected by platform (P < 0.001). Source quality differed by platform (P = 0.037).
Conclusion:
AI platforms could be resources; currently they fall short of the knowledge that Equine Extension Specialists can offer. AI platforms had difficulty addressing complex topics and demonstrated inconsistent performance across criteria.
