临床决策支持中的人类和大型语言模型:使用医学计算器进行的一项研究
Nicholas C Wan1, Qiao Jin1, Joey Chan1
1Division of Intramural Research, National Library of Medicine (NLM), National Institutes of Health (NIH), Bethesda, MD, USA.
大型语言模型 (LLM) 在推医疗计算器方面显示出有限的准确性,不足以达到人类的性能. 了解和知识差距阻碍了他们的临床决策支持能力.
科学领域:
- 人工智能在医学中的应用
- 临床决策支持系统 临床决策支持系统
- 医疗信息学 医疗信息学
背景情况:
- 大型语言模型 (LLM) 越来越多地被评估为医学知识,但它们在临床决策中的实用性,特别是用于工具选择,尚不清楚.
- 评估LLM推适当医疗计算器的能力对于安全有效的临床实践至关重要.
研究的目的:
- 评估各种大型语言模型 (LLM) 在推医疗计算器方面的表现,与人类表现相比.
- 为了确定LLM在选择临床计算器时会犯的错误类型.
主要方法:
- 九个LLM (开源,专有,域特定) 在1009个选择题中进行了测试,涵盖35个临床计算器.
- 在100个问题的子集上,LLM的表现与人类注释者进行了比较.
- 对表现最高的法学士进行了错误分析.
主要成果:
- 性能最好的LLM在回答有关医疗计算器的问题时达到66.0%的准确率.
- 人类注释者平均准确率为79.5%,超过了LLM.
- 在LLM的错误主要是由于理解 (49.3%) 和计算器知识缺陷 (7.1%).
结论:
- 当前的大型语言模型 (LLM) 在推医疗计算器的临床决策方面并不超过人类的性能.
- 在临床计算器的选择中,LLM需要在理解和医学知识方面显著改进,才能成为可靠的临床计算器选择工具.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:51Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
相关概念视频
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Fundamental Mathematical Principles in Pharmacokinetics: Calculus and Graphs
On the other hand, integral calculus focuses on...
Mathematical Modeling: Problem Solving
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Statistical Software for Data Analysis and Clinical Trials
