臨床意思決定支援における人間と大規模言語モデル:医療計算機を用いた研究
Nicholas C Wan1, Qiao Jin1, Joey Chan1
1Division of Intramural Research, National Library of Medicine (NLM), National Institutes of Health (NIH), Bethesda, MD, USA.
大規模言語モデル(LLM)は、医療計算機の推奨において限定的な精度しか示さず、人間のパフォーマンスに及ばない。理解力と知識のギャップが、臨床意思決定支援能力を妨げている。
科学分野:
- 人工知能と医学
- 臨床意思決定支援システム
- 医療情報学
背景:
- 大規模言語モデル(LLM)は医療知識に関してますます評価されていますが、臨床意思決定、特にツールの選択における有用性は不明です。
- 適切な医療計算機を推奨するLLMの能力を評価することは、安全で効果的な臨床実践のために不可欠です。
研究 の 目的:
- 人間のパフォーマンスと比較して、医療計算機を推奨する際の様々な大規模言語モデル(LLM)のパフォーマンスを評価すること。
- 臨床計算機を選択する際にLLMが犯すエラーの種類を特定すること。
主な方法:
- 9つのLLM(オープンソース、プロプライエタリ、ドメイン固有)を、35の臨床計算機を対象とした1,009の多肢選択問題でテストしました。
- LLMのパフォーマンスを、100問のサブセットで人間のアノテーターと比較しました。
- 最もパフォーマンスの高いLLMについてエラー分析を実施しました。
主要な成果:
- 最もパフォーマンスの高いLLMは、医療計算機に関する質問に66.0%の精度で回答しました。
- 人間のアノテーターは平均79.5%の精度を達成し、LLMを上回りました。
- LLMのエラーは、主に理解力(49.3%)と計算機の知識不足(7.1%)によるものでした。
結論:
- 現在のLLMは、臨床意思決定のための医療計算機の推奨において、人間のパフォーマンスを上回るものではありません。
- LLMが臨床計算機の選択のための信頼できるツールとなるためには、理解力と医学知識の大幅な向上が必要です。
さらに関連する動画
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:51Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
関連する概念動画
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Fundamental Mathematical Principles in Pharmacokinetics: Calculus and Graphs
On the other hand, integral calculus focuses on...
Mathematical Modeling: Problem Solving
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Statistical Software for Data Analysis and Clinical Trials
