评估人工智能与人类创建的MCQ在儿科医学教育中的实用性
James Knight1,2, Richard G McGee2,3,4, Bunmi S Malau-Aduli1,2
1Faculty of Medicine and Health, School of Rural Health, The University of New England, Armidale, NSW, Australia.
Journal of medical education and curricular development
|March 6, 2026
概括
生成型人工智能 (AI) 显示出创造医学教育多选择题 (MCQ) 的潜力,但人类撰写的问题目前在儿科评估中显示出优越的质量和一致性.
科学领域:
- 医学教育 医学教育
- 人工智能的人工智能
- 心理测量 心理测量 心理测量
背景情况:
- 多选题 (MCQ) 对于医学教育评估至关重要,但其创建需要大量资源.
- 生成型人工智能 (AI) 为自动化MCQ开发提供了一个潜在的解决方案.
- 人工智能生成的MCQ的心理测量质量与人类撰写的项目相比,特别是在儿科教育中,尚不清楚.
研究的目的:
- 直接比较人工智能生成和人类生成的儿科MCQ的质量.
- 通过基于经典测试理论的项目分析来评估心理测量属性.
主要方法:
- 用AI (Microsoft Copilot) 和人类生成的儿科MCQ进行了形成性考试,对4年级医学学生进行了考试.
- 项目分析计算了难度和歧视指数,项目总和相关性,以及分心器的功能.
- 使用KR-20评估可靠性,并进行对联t测试,比较AI和人类项目性能.
主要成果:
- 人类撰写的MCQ在所有质量指标上都超过了人工智能生成的问题.
- 人工智能问题具有较低的歧视性 (0.19对0.29),并且在可接受的难度范围之外的比例更高 (56%对32%).
- 干扰因子分析偏好了人类问题,不起作用的干扰因子较少,人工智能项目质量更一致.
结论:
- 目前的生成人工智能不能始终产生高质量的儿科MCQ,与人类专业知识相提并论.
- 人工智能作为医疗教育评估的混合人类-人工智能工作流程中的补充工具显示出希望.
- 在将AI整合到评估设计中,平衡可靠性,有效性,可接受性和成本效益是关键.
相关概念视频
Methods of Documentation III: PIE
2.1K
Problem-intervention-evaluation (PIE) is a systematic approach to documentation used in healthcare settings for clinical decision-making and patient care planning. It is a structured approach to organizing patient data based on problems, interventions, and evaluations. Here's a breakdown of its key features and considerations:
2.1K
The Availability Heuristic
7.2K
A heuristic is a general problem-solving framework (Tversky & Kahneman, 1974). You can think of these as mental shortcuts that are used to solve problems. Different types of heuristics are used in different types of situations, and the impulse to use a heuristic occurs when one of five conditions is met (Pratkanis, 1989):
7.2K
