确定人工智能程序在医学院董事会考试中的可靠性
1Arif Keskin, Department of Anatomy, Faculty of Medicine, Giresun University, 28100, Giresun, Turkiye.
Pakistan journal of medical sciences
|December 26, 2025
概括
像Copilot,ChatGPT和Gemini这样的生成人工智能 (AI) 工具在解剖学考试中显示出比医学学生更高的准确性. 然而,由于临床推理的局限性,人工智能应该补充,而不是取代传统教育.
科学领域:
- 医学教育 医学教育
- 人工智能的人工智能
- 人体解剖学 解剖学 解剖学
背景情况:
- 生成型人工智能应用越来越多地用于医学教育.
- 在评估医学学生的知识时,评估AI可靠性至关重要.
研究的目的:
- 评估生成AI程序 (微软Copilot,谷歌Gemini,OpenAI ChatGPT) 与一年级和二年级医学学生在解剖学考试中的表现的准确性.
- 确定AI工具在基础医学科学教育中的可靠性.
主要方法:
- 从2023-2024学年使用了286个解剖学问题.
- 根据学生的表现,根据难度分类问题.
- 人工智能工具回答了同样的问题提出给学生.
主要成果:
- 微软的Copilot (97.7%),OpenAI的ChatGPT (94.4%) 和谷歌的Gemini (86.5%) 的准确性高于学生.
- 双子座和ChatGPT在非常困难的问题上表现与学生相似.
- 双子座的表现各不相同,在基本知识问题上的临床解释方面表现出色.
结论:
- 虽然人工智能工具超过了学生的准确性,但它们不适合在解剖学教育中无监督使用.
- 人工智能缺乏临床推理和人类经验,因此需要将其用作受控的补充教育工具.
相关概念视频
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K
Issues And Trends In Healthcare Delivery System
6.1K
The issues and trends in healthcare delivery are constantly changing. The COVID-19 pandemic is one recent issue that wreaked havoc on healthcare systems, causing a shortage of healthcare workers, high demand for medicines and supplies, and increased medical expenditure due to a lack of insurance. Other issues include rising healthcare costs and care fragmentation.
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
6.1K

