人工智能生成的单个最佳答案问题的质量保证和有效性
Ayla Ahmed1, Ellen Kerr1, Andrew O'Malley2
1University of St Andrews, St Andrews, UK.
BMC medical education
|February 26, 2025
概括
生成型人工智能 (AI) 可以创建医学教育考试问题,解决评估银行短缺问题. 人工智能生成的问题与人类撰写的问题相比,性能相对较好,尽管质量检查至关重要.
科学领域:
- 医学教育 医学教育
- 人工智能的人工智能
- 评估评估的方法
背景情况:
- 生成型人工智能的进步提供了新的教育工具,特别是在医疗领域.
- 本研究探讨了评估问题库的枯竭以及对更具形成性的评估的需求.
- 它探讨了人工智能产生考试问题的潜力,而不仅仅是通过现有的考试.
研究的目的:
- 评估生成性AI在为医学教育创建单一最佳答案 (SBA) 问题方面的实用性.
- 评估人工智能生成问题的质量和性能,与人类撰写的问题相比.
- 确定人工智能是否可以帮助补充和多样化医疗评估资源.
主要方法:
- 使用OpenAI GPT-4生成了220个SBA问题,与ScotGEM学习成果保持一致.
- 一个专家小组审查了问题的准确性和质量;69%是适合使用的.
- 一年级和二年级学生进行了两次50项 (25项人工智能,25项人类) 的培训考试.
主要成果:
- 69%的人工智能生成的SBA问题需要轻微编辑才能包含;31%由于不准确或不整齐而被拒绝.
- 在人工智能生成的问题和人类撰写的问题之间,在性能 (便利性和歧视指数) 中没有发现显著差异.
- 人工智能产生的问题遵循医学院理事会评估联盟的指导方针.
结论:
- 人工智能大语言模型 (LLM) 可以生成高质量的SBA问题,与教育指南和学习成果保持一致.
- 严格的质量保证过程对于识别和拒绝不准确的AI产生的问题至关重要.
- 终身医学课程显示了补充传统问题生成的巨大潜力,增强了医学评估资源.
更多相关视频
相关概念视频
Quality Assurance
112
Quality assurance is the overarching term used to describe the activities employed to ensure the proper performance of a system. These activities can be classified into three categories: quality control, quality assessment, and internal corrective measures. Typically, these activities work cyclically: quality control is performed before and during the analysis, while quality assessment occurs during and after the investigation. Internal corrective measures are implemented based on the findings...
112
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Data Validation
4.9K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
4.9K
Accuracy and Errors in Hypothesis Testing
169
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
169
Quality Control
145
Quality control is one of the three cyclical quality assurance activities that help keep a system under statistical control. Typical quality control activities include creating quality control charts, conducting proficiency testing, and documenting and archiving results.
Quality control helps track data, visualize trends, and identify variations, making it easier to detect deviations that may affect the accuracy of an analysis. One way to do this is by generating a quality control chart, which...
Quality control helps track data, visualize trends, and identify variations, making it easier to detect deviations that may affect the accuracy of an analysis. One way to do this is by generating a quality control chart, which...
145
Measures of Intelligence
6.0K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
6.0K


