医学的な短い答えの質問の格付け者としての大規模な言語モデルの評価:専門家の人間の格付け者との比較分析
Olena Bolgova1, Paul Ganguly1, Muhammad Faisal Ikram1
1College of Medicine, Alfaisal University, Riyadh, Kingdom of Saudi Arabia.
Medical education online
|August 24, 2025
まとめ
大型言語モデル (LLM) は医学的な短答えの質問の格付けに有望ですが,性能はモデルと医学分野によって異なります. 専門家の基準は一貫してLLMの精度を向上させず,ドメイン特有の実装と人間の監督を必要とした.
科学分野:
- 医療教育の評価
- 医療における人工知能
- 自然言語処理
背景:
- 短回答質問 (SAQ) の医療教育評価は時間がかかり,専門家の意見が必要です.
- 大型言語モデル (LLM) は,SAQの格付けを自動化するための潜在的な解決策です.
- 専門的な医療評価の文脈におけるLLMの有効性は十分に確立されていません.
研究 の 目的:
- 5つのLLMの格付けパフォーマンスを,医療SAQの専門家人間格レーダーと比較する.
- 異なる医学分野 (解剖学,組織学,胚学,生理学) に関するLLM能力を比較する.
- 専門家によって定義された格付けラベルの提供がLLMのパフォーマンスに与える影響を評価する.
主な方法:
- 4つの医療科目の804人の学生のSAQ応答の分析
- 3人の専門家による回答の評価
- 5つのLLM (GPT-4.1,Gemini,Claude,Copilot,DeepSeek) の評価は,2つのアプローチを用いて行われました.
- コヘンのカッパとクラス内相関係数 (ICC) を用いた合意の測定.
主要な成果:
- 専門家の間でかなりの合意が見られた (平均カッパ: 0. 69,ICC: 0. 86).
- LLMの成績は質問の種類やモデルによって大きく異なっており,どのLLMも一貫して他のLLMを上回らない.
- 専門家とLLMの間で最も高い合意は,1つの質問でClaude (カッパ:0.61) と別の質問でDeepSeek (カッパ:0.53) が達成した.
- 専門家の基準がLLMの業績に与える影響は不一致でした.
- 専門家に比べてLLMの格付けの厳しさは大きく異なっており,分野別格付けの違いもありました.
結論:
- LLMは,医療のSAQ評価をサポートする可能性を証明しているが,ドメイン特有の性能の変動を示している.
- 専門家ラブリックの提供は,LLMの格付けの精度を一貫して向上させなかった.
- 医学教育評価におけるLLMの成功には,特定の領域を慎重に検討し,継続的な人間の監督が必要です.
さらに関連する動画
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
681
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
575
関連する概念動画
Classification of Illness
9.4K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
9.4K
Sensitivity, Specificity, and Predicted Value
1.9K
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
1.9K
