在基于病例的牙科诊断中,Manus,ChatGPT和Claude的准确性和可靠性
Ahmed A Madfa1, Abdullah F Alshammari2, Bassam A Anazi3
1Department of Restorative Dental Science, College of Dentistry, University of Ha'il, Ha'il, Saudi Arabia.
Frontiers in oral health
|January 26, 2026
概括
克劳德和马努斯在牙科场景中显示出比ChatGPT更好的诊断准确性和一致性. 虽然有希望,但这些人工智能模型需要进一步大规模评估以进行临床整合.
科学领域:
- 人工智能在牙科中的应用
- 医疗保健 教育 技术 技术
- 临床决策支持系统 临床决策支持系统
背景情况:
- 人工智能 (AI),特别是大型语言模型 (LLM),越来越多地用于医疗保健教育和临床决策.
- 现有的LLM如ChatGPT和Claude在医疗领域显示出潜力,但它们在牙科的诊断能力尚未得到充分研究.
- 像Manus这样的新兴AI平台需要对它们在牙科诊断中的实用性进行评估.
研究的目的:
- 用现实的,基于案例的牙科场景来比较ChatGPT,Claude和Manus的诊断准确性和一致性.
- 在牙科诊断环境中评估不同AI模型的性能.
主要方法:
- 一个标准化测试,涉及117个多选题从牙科小贴士给了ChatGPT,克劳德和马努斯在两个时间点.
- 答案与专家验证的答案密钥进行了评估.
- 统计分析包括科恩的卡帕对评价者之间的可靠性和奇平方,麦克内马尔和t测试进行比较.
主要成果:
- 与ChatGPT相比,克劳德和曼努斯在两个测试阶段都表现出卓越的诊断准确性和一致性.
- 在第二轮中,Claude和Manus的准确率达到92.3%,而ChatGPT的准确率为76.9%.
- 克劳德和马努斯的模型内部一致性 (科恩的卡帕 = 0.714和0.782) 比ChatGPT (卡帕 = 0.560) 高,尽管差异在统计学上并不显著.
结论:
- 克劳德和曼努斯的诊断性能和响应稳定性比ChatGPT数字上更高,但这些差异缺乏统计学意义.
- 观察到的变异性需要进行更大规模的研究来证实发现,并指导用于牙科实践和教育的AI工具选择.
- 在将AI整合到牙科中时,考虑准确性和一致性至关重要.
相关概念视频
Reliability and Validity
13.8K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
13.8K
Improving Translational Accuracy
14.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.9K
Improving Translational Accuracy
3.7K
3.7K
Uncertainty in Measurement: Accuracy and Precision
100.8K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
100.8K
Accuracy and Precision
15.1K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value. Highly accurate...
15.1K
Accuracy, limits, and approximation
1.3K
Accuracy, limits, and approximations are common in many fields, especially in engineering calculations. These concepts are imperative for ensuring that a given value is as close as possible to its true value.
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
1.3K


