在儿科牙科学术GPT和ChatGPT-4Omni的比较评估:准确性和完整性的分析
Alev Eda Okutan1, Berkant Sezer2
1Pediatric dentist in private practice, Istanbul.
Pediatric dentistry
|January 19, 2026
概括
与ChatGPT-4 Omni相比,ScholarGPT在回答儿科牙科临床问题方面表现出更高的准确性和完整性. 这凸显了专用人工智能工具在牙科教育和实践中的潜力.
科学领域:
- 人工智能在牙科中的应用
- 临床决策支持系统 临床决策支持系统
- 牙科教育技术 牙科教育技术
背景情况:
- 人工智能 (AI) 的快速发展为支持临床决策和儿童牙科等专业领域的教育提供了机会.
- 像ChatGPT-4 Omni (ChatGPT-4o) 这样的大型语言模型 (LLM) 和像ScholarGPT这样的特定领域工具正在越来越多地被探索在医疗保健中的实用性.
研究的目的:
- 批判性地评估和比较ScholarGPT和ChatGPT-4o对儿科牙科临床问题的答案的准确性和完整性.
- 评估这些人工智能工具在儿科牙科内的各种子专业主题的性能.
主要方法:
- 制定了30个临床问题,涵盖了6个儿科牙科主题.
- 来自ScholarGPT和ChatGPT-4o的回复被六位经验丰富的儿科牙医收集并独立评估.
- 评估人员使用利克特尺度来评估答案的准确性 (事实正确性,相关性,连贯性) 和完整性,并使用非参数测试进行统计分析.
主要成果:
- 与ChatGPT-4o (准确度:4;完整度:2) 相比,ScholarGPT在所有主题 (P<0.001) 中取得了显著更高的中位准确度 (5分) 和完整度 (3分).
- 两种AI模型都显示了特定主题的精度差异,ScholarGPT在"裂密封剂"方面表现出色,ChatGPT-4o在"化物"方面表现出色.
- 在ScholarGPT的各个主题中,完整度得分有很大差异,但在ChatGPT-4o中却没有.
结论:
- 在提供对儿科牙科临床问题的准确和完整答案方面,ScholarGPT的表现优于ChatGPT-4o.
- 特定领域的人工智能工具显示出增强牙科教育和提供临床支持的希望,尽管需要进一步开发和验证.
相关概念视频
Comparing the Survival Analysis of Two or More Groups
570
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
570
Uncertainty in Measurement: Accuracy and Precision
100.1K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
100.1K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.6K
3.6K
Accuracy and Precision
14.2K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value. Highly accurate...
14.2K
Accuracy, limits, and approximation
1.1K
Accuracy, limits, and approximations are common in many fields, especially in engineering calculations. These concepts are imperative for ensuring that a given value is as close as possible to its true value.
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
1.1K


