从人类和人工智能 (GPT-4和Gemini) 的角度评估和比较学生在考试中的反应
Kubra Yildiz Domanic1, Sukran Baycan2
1Atlas University, Istanbul, Turkey. drkubrayildiz@gmail.com.
BMC medical education
|October 2, 2025
概括
像ChatGPT和Gemini这样的生成人工智能 (AI) 模型显示出牙科教育的潜力. 与Gemini相比,ChatGPT在回答牙科检查问题方面表现出更高的准确性和一致性.
科学领域:
- 牙科教育 牙科教育
- 教育中的人工智能
- 健康 职业 教育 教育 专业
背景情况:
- 包括ChatGPT (GPT-4) 和Gemini在内的生成人工智能模型为提高牙科教育提供了机会.
- 这些人工智能工具在改善牙科假体技术 (DPT) 和口腔健康 (OH) 计划中的学习和评估方面表现有前途.
研究的目的:
- 评估GPT-4和Gemini在回答牙科教育考试问题的准确性,可靠性和一致性.
- 该研究的重点是多选项,真假和短答案问题格式.
主要方法:
- 一项探索性研究使用了来自DPT和OH课程的30个问题 (10个MCQ,10个T/F,10个SAQ).
- 聊天GPT和Gemini两次回答问题,以评估响应的一致性,由两个独立研究人员对预定义的关键进行评估.
- 数据分析包括描述性统计,卡帕一致系数和千平方测试.
主要成果:
- 聊天GPT在MCQ (90%) 和T/F (85%) 中取得了高精度,在SAQ (60%) 中表现较低.
- 双子座的准确度在60%至70%之间,在SAQ中准确度最高 (70%).
- 聊天GPT表现出显著的一致性 (Kappa=0.754,p=0.001),而双子座表现出较少的一致性 (Kappa=0.634,p=0.001).
结论:
- 与Gemini相比,ChatGPT在结构化牙科教育评估中表现出更高的准确性和一致性.
- 人工智能工具可以增强牙科教学和评估策略,支持个性化的学习和学术完整性,当被深思熟虑地整合时.
相关概念视频
Comparing Experimental Results: Student's t-Test
4.9K
The t-test is a statistical method used to compare the sample mean with a population mean or compare two means from two data sets. The test statistic is calculated from the standard deviation, mean, and number of measurements in the data set at a selected confidence interval and then compared to a table of critical values at this confidence level. If the test statistic is smaller than the critical value, the null hypothesis is accepted. In this case, we state that the difference between the...
4.9K
Non-equilibrium in the Cell
5.3K
An important concept in studying metabolism and energy is that of chemical equilibrium. Most chemical reactions are reversible. They can proceed in both directions, releasing energy into their environment in one direction, and absorbing it from the environment in the other direction. The same is true for the chemical reactions involved in cell metabolism, such as the breaking down and building up of proteins into and from individual amino acids, respectively. Reactants within a closed system...
5.3K
Measures of Intelligence
8.3K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
8.3K
Intelligence
8.4K
The term "intelligence" is complex because it refers to both behavior and individuals, and its interpretation varies across cultures. European Americans tend to link intelligence with reasoning and cognitive skills, while in Kenya, it is tied to responsible participation in family and social life. In Uganda, intelligence is seen as the ability to know the right actions and carry them out effectively, while the Iatmul people of Papua New Guinea associate it with the capacity to remember...
8.4K
Multiple Comparison Tests
4.4K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.4K
Reliability and Validity
13.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.7K


