新しく開発された大規模な言語モデルの診断性能:重症症例における比較研究
Xintong Wu1, Yu Huang1, Qing He1
1Department of Intensive Care Medicine, Affiliated Hospital of Southwest Jiaotong University, The Third People's Hospital of Chengdu, Chengdu, Sichuan, China.
International journal of medical informatics
|August 27, 2025
まとめ
新しく開発された大型言語モデル (LLM) は,集中治療室 (ICU) での臨床的意思決定支援に希望を示しています. ChatGPT-o3は診断の精度でリードし,オープンソースのDeepSeek-R1も競争力を持っています.
科学分野:
- 医療における人工知能
- 臨床的意思決定支援システム
- クリティカル ケア 医療
背景:
- 大規模言語モデル (LLM) は,臨床意思決定のサポートの可能性を示しています.
- 集中治療室 (ICU) でのLLM診断能力は十分に研究されていない.
- この研究は,重症病の診断のために新たに開発されたLLMを評価します.
研究 の 目的:
- 4つの新しいLLMの診断の正確性を評価する.
- これらのLLMの差異診断の質と応答の質を比較する.
- 集中治療室での臨床決定支援に最も有効なLLMを特定する.
主な方法:
- 横断的な比較研究では,ChatGPT-4o,ChatGPT-o3,DeepSeek-V3,およびDeepSeek-R1の4つのLLMが評価されました.
- LLMは,ICUの出版物から50件の重症症例でテストされました.
- 診断の精度,差異診断の質,応答の質を比較した.
主要な成果:
- ChatGPT-o3は診断の精度が最も高い (72%),次にDeepSeek-R1 (68%) とChatGPT-4o (64%) が続いた. ディープシークV3の精度は32%でした
- ChatGPT-o3,DeepSeek-R1,およびChatGPT-4oはDeepSeek-V3を大幅に上回った.
- すべての評価されたLLMは高い応答品質 (完全性,明瞭性,有用性) を示し,ChatGPT- o3とDeepSeek- R1は優れた差異診断品質を示した.
結論:
- 新しく開発されたLLM,特に推論モデルは,重症治療における診断をサポートする大きな可能性を示している.
- ドメイン固有の微調整は,LLMの診断精度をさらに高めることができます.
- オープンソースのDeepSeek-R1の競争力は,リソースが限られた環境での可能性を強調しています.
関連する概念動画
Classification of Illness
7.9K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
7.9K
Language and Cognition
438
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
438
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K


