大型语言模型在兽医本科多选择题考试中的表现:比较评估
Santiago Alonso Sousa1, Syed Saad Ul Hassan Bukhari1, Paulo Vinicius Steagall1,2
1Department of Veterinary Clinical Sciences, Jockey Club College of Veterinary Medicine and Life Sciences, City University of Hong Kong, Kowloon, Hong Kong SAR, China.
Frontiers in veterinary science
|September 11, 2025
概括
先进的人工智能,特别是大型语言模型 (LLM),在兽医教育中显示出强大的潜力. 聊天GPT模型在兽医考试中表现出卓越的表现,突出了LLM作为有价值的评估工具.
科学领域:
- 兽医医学 兽医医学 兽医医学
- 人工智能的人工智能
- 教育技术的教育技术
背景情况:
- 人工智能 (AI) 的应用,特别是大型语言模型 (LLM),在兽医教育和实践中正在出现.
- 然而,它们在专门的兽医环境中的有效性需要进一步研究.
研究的目的:
- 在兽医多选题 (MCQ) 上对九个高级LLM进行比较性绩效评估.
- 在兽医评估中确定影响LLM绩效的因素.
主要方法:
- 九个LLM在250个兽医本科期末考试的MCQ上进行了测试.
- 问题涵盖了各种物种,临床主题,推理阶段,并包括基于文本和图像的格式.
主要成果:
- 聊天GPT模型 (o1Pro和4.5) 的精度最高 (90.4%和90.8%),Kimi 1.5 的精度最低 (64.8%).
- 性能随着问题难度的下降而下降,并且对于基于图像的问题来说更低.
- 开放AI模型显示了增强的视觉解释能力.
结论:
- 作为对兽医评估设计质量保证的支持工具,LLM显著有前途.
- 问题难度,格式和特定领域的培训数据是影响LLM绩效的关键因素.
相关概念视频
Multiple Comparison Tests
4.4K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.4K
Improving Translational Accuracy
3.6K
3.6K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K

