七个大语言模型在解剖学考试问题上的表现
Weronika Chaba-Karnaś1, Natalia Kozioł1, Natalia Kalita1
1Department of Anatomy, Jagiellonian University Medical College, Kraków, Poland.
概括
人工智能 (AI) 模型在解决医学解剖学考试方面表现出很高的效率,表现最好的模型的准确率超过80%. 这种能力引起了教育工作者对学生评估和学术完整性的担忧,因为人工智能技术在不断发展.
科学领域:
- 医疗教育 医学教育
- 人工智能的人工智能
- 人体解剖学 解剖学 解剖学
背景情况:
- 人工智能 (AI) 和大型语言模型 (LLM) 的快速发展需要评估它们在医学等专业领域的实用性.
- 人工智能对医学教育的潜在影响,特别是在解剖学评估中,需要进行彻底的调查.
研究的目的:
- 评估各种AI语言模型在回答专为医学学生设计的理论解剖学考试问题的有效性.
- 将人工智能性能与人类学生在解剖学知识方面的能力进行基准测试.
主要方法:
- 从过去的医学学生考试中利用了555个多选择解剖学问题 (150个波兰语,405个英语).
- 测试了七个人工智能模型:ChatGPT-4o mini,ChatGPT-4o,DeepSeek,Copilot,Gemini,Bielik和PLLum. 这些都是我们测试的.
- 分析基于问题类型和解剖学主题的AI性能.
主要成果:
- 聊天GPT-4o获得了最高的准确度 (83.1%),其次是Copilot (79.6%) 和Gemini (78.8%).
- 人工智能模型通常在有关器官功能的问题上表现更好 (75%),在多个答案问题上表现更差 (37.6%).
- 大多数经过测试的AI模型都证明了能够独立通过解剖学考试的能力.
结论:
- 目前的AI语言模型在回答解剖学考试问题方面具有显著的能力,可能超过人类学生的通过率.
- 教师必须调整评估策略以考虑AI能力,因为AI性能正在迅速提高.
- 开发抗人工智能评估方法对于保持医学教育学术完整性至关重要.
更多相关视频
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
3.2K
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
相关概念视频
Classification of Bones
9.5K
The bones of the human skeletal system are of varied shapes, sizes, and functions. They can be classified based on their shape and function into four major classes: long bones, short bones, flat bones, and irregular bones. Some classifications include a fifth type, the sesamoid bones, as a separate class, whereas others categorize them under short bones.
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
9.5K
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
