人工智能与人类生成的医学教育多选择题对比:高风险考试中的队列研究
Alex Kk Law1,2, Jerome So3, Chun Tat Lui4
1The Accident and Emergency Medicine Academic Unit (AEMAU), The Chinese University of Hong Kong (CUHK), 2nd Floor, Main Clinical Block and Trauma Centre, Prince of Wales Hospital, Shatin, Hong Kong, China. alexlaw@cuhk.edu.hk.
BMC medical education
|February 8, 2025
概括
大型语言模型 (LLM) 可以有效地生成医学教育的多选择题 (MCQ),但人类专家审查对于确保高风险评估的准确性和深度至关重要.
科学领域:
- 医疗教育 技术 技术 医学教育
- 医疗保健中的人工智能
- 评估和评价的评估和评估.
背景情况:
- 高质量的多选择题 (MCQ) 对于医学教育评估至关重要.
- 传统的MCQ创建是资源密集型和耗时的.
- 大型语言模型 (LLM) 为高效的MCQ生成提供了一个潜在的解决方案.
研究的目的:
- 评估ChatGPT-4o生成的MCQ的质量和心理特征.
- 在高风险的医疗执照考试中,将人工智能生成的MCQ与人类创建的MCQ进行比较.
- 评估在医学教育中使用LLM用于可扩展的MCQ创建的可行性.
主要方法:
- 一项前性队列研究,涉及准备进行急诊医学初级检查 (PEEM) 的医生.
- 参与者完成了两组100个MCQ:一个是人工智能生成的,一个是人类生成的.
- 专家评审人员评估了MCQ的正确性,相关性,难度,布鲁姆的分类法对齐和项目写作缺陷,以及心理测量分析.
主要成果:
- 人工智能生成的MCQ更容易,但与人类MCQ相比,其歧视指数类似.
- 专家审查显示,人工智能生成的MCQ中存在更多事实上的不准确性和不相关性.
- 人工智能问题主要测试低级认知技能,而人类MCQ更好地评估高级技能,尽管人工智能显著减少了问题生成时间.
结论:
- 聊天GPT-4o显示了有效的MCQ生成的潜力,但需要专家监督高风险的医疗评估.
- 将人工智能效率与人类专业知识相结合,可以优化医疗教育的MCQ创建.
- 这种方法提供了一个可扩展的模型,平衡评估的时间效率和内容质量.
相关概念视频
Surveys
14.7K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
14.7K
Multiple Comparison Tests
3.8K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.8K
Cochran's Q Test
197
Cochran's Q Test is a nonparametric statistical test used to determine if there are potential differences in the outcomes of three or more related groups on a binary (yes/no) or dichotomous outcome. It is essentially an extension of the McNemar Test, which is limited to two related samples - Cochran's Q test can handle three or more related samples, making it more versatile in scenarios where subjects are measured under multiple conditions. The test statistic follows a Chi-Square...
197
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
114
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
114
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Blind Procedures
10.6K
Ideally, the people who observe and record the children’s behavior are unaware of who was assigned to the experimental or control group, in order to control for experimenter bias. Experimenter bias refers to the possibility that a researcher’s expectations might skew the results of the study. Remember, conducting an experiment requires a lot of planning, and the people involved in the research project have a vested interest in supporting their hypotheses. If the observers knew which...
10.6K


