併合性疾患患者の診断のための大規模な言語モデルの急速なベンチマーク:LLM-as-a-Judge方法を活用した比較研究
Peter Sarvari1, Zaid Al-Fagih1
1Rhazes AI, First Floor, 85 Great Portland Street, London, W1W 7LT, United Kingdom.
JMIRx med
|August 29, 2025
まとめ
Gemini 2.5は,実際の患者データで評価された21の大型言語モデル (LLM) の間で優れた診断精度を示した. この研究は,医療診断の改善におけるLLMの可能性を強調しているが,さらなる研究が必要である.
科学分野:
- 医療における人工知能
- 臨床的意思決定支援システム
- 医療のための自然言語処理
背景:
- 診断誤差は患者の死亡率に大きく寄与し,米国で3番目に多い死因となっています.
- 大規模言語モデル (LLM) は臨床医の診断を支援する見込みが高く,実際の患者コホートに関する比較性能データは不足しています.
研究 の 目的:
- 大規模で実在する患者のデータセットを用いた18の一般的な大型言語モデル (LLM) の診断能力を比較する.
- 異なるプロンプトと温度設定がLLMの診断性能に与える影響を評価する.
- LLMの診断精度を高めるためのリトリーバルの拡張生成 (RAG) の有効性を評価する.
主な方法:
- ランダムに選択された1000人の医療情報マートで21人のLLMを評価しました.
- 自動評価のためのLLM-as-a-judgeアプローチを採用し,患者の記録からの最終的な診断コードと比較した.
- 診断のヒット率を計算し,比率の集合 z テストを使用して統計的有意性を評価した.
主要な成果:
- Gemini 2.5は,GPT-4.1によって評価されたとき,GPT-4.1やClaude-4 Opusのような他の主要なモデルを上回る最も高い診断ヒット率 (97.4%) を達成しました.
- GPT-4.1は,GPT-4 Turboによる別々の評価で最高のパフォーマンスを示し,審査員LLMによって評価の変動を示しました.
- 検索強化生成 (RAG) は,GPT-4o 05-13のヒット率を0. 8% (P<. 006) で大幅に改善し,性能は異なるプロンプトによって変化した.
結論:
- LLMは臨床での診断の精度を向上させる大きな可能性を示しています.
- LLMの診断ツールを完全に理解し,実装するには,多様なデータセットと臨床検証によるさらなる研究が不可欠です.
- 人工知能開発者と医師の緊密な協力は,医療におけるLLMの責任ある統合に不可欠です.
さらに関連する動画
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
9.3K
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
570
関連する概念動画
Classification of Illness
7.9K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
7.9K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
708
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
708
Mechanistic Models: Compartment Models in Individual and Population Analysis
85
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
85
