大型语言模型在心血管认证模拟考试中的比较性能.
Eesha Nachnani1, Kashish Goel2, Alexander E Sullivan2
1University School of Nashville, Nashville, Tennessee, USA.
American heart journal
|January 14, 2026
概括
人工智能 (AI) 越来越多地用于医学. 一项研究发现,在测试的AI模型中,只有ChatGPT-4.0在心血管医学考试中表现得与人类相比.
科学领域:
- 医疗信息学 医疗信息学
- 医疗保健中的人工智能
- 心血管医学 心血管医学
背景情况:
- 人工智能 (AI) 在医疗实践中展示了越来越大的能力.
- 人工智能工具正在评估它们在临床决策和检查方面的潜在帮助.
研究的目的:
- 评估领先的大型语言模型 (LLM) 在心血管医学委员会样式检查中的表现.
- 为了比较不同AI平台在专业医疗领域的有效性.
主要方法:
- 评估了三种受欢迎的LLM (ChatGPT-4.0,Gemini,Bing AI) 的使用情况.
- 性能与心血管医学委员会样式的考试进行了测量,与人类参与者相比得分更高.
主要成果:
- 在心血管检查中,ChatGPT-4.0的得分与人类参与者相当.
- 双子座和Bing AI在这个专业医疗评估中没有达到类似的性能水平.
结论:
- 聊天GPT-4.0显示了医学知识评估在心脏病学显著的潜力.
- 需要进一步的研究来评估人工智能在专业医疗领域的更广泛的临床应用性和局限性.
相关概念视频
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K

