解剖学試験問題に対する7つの大規模言語モデルのパフォーマンス
Weronika Chaba-Karnaś1, Natalia Kozioł1, Natalia Kalita1
1Department of Anatomy, Jagiellonian University Medical College, Kraków, Poland.
まとめ
人工知能(AI)モデルは医学解剖学試験の解答において高い有効性を示しており、トップパフォーマーは正答率80%を超えています。この能力は、AI技術の進歩に直面した学生の評価と学業の完全性に関して、教育者に懸念を抱かせています。
科学分野:
- 医学教育
- 人工知能
- 解剖学
背景:
- 人工知能(AI)と大規模言語モデル(LLM)の急速な進歩は、医学のような専門分野におけるそれらの有用性を評価する必要性を生じさせています。
- 医学教育、特に解剖学評価におけるAIの潜在的な影響は、徹底的な調査が必要です。
研究 の 目的:
- 医学部学生向けの理論的な解剖学試験問題に答える上での様々なAI言語モデルの有効性を評価すること。
- AIのパフォーマンスを解剖学の知識における人間の学生の能力と比較してベンチマークすること。
主な方法:
- 過去の医学部学生の試験からの555の多肢選択式解剖学問題(ポーランド語150問、英語405問)を使用しました。
- ChatGPT-4o mini、ChatGPT-4o、DeepSeek、Copilot、Gemini、Bielik、PLLumの7つのAIモデルをテストしました。
- 質問の種類と解剖学の主題に基づいてAIのパフォーマンスを分析しました。
主要な成果:
- ChatGPT-4oは最高の精度(83.1%)を達成し、次いでCopilot(79.6%)、Gemini(78.8%)が続きました。
- AIモデルは一般的に臓器機能(75%)に関する質問でより良い成績を収め、複数回答の質問(37.6%)では成績が悪かったです。
- テストされたほとんどのAIモデルは、解剖学試験に独立して合格する能力を示しました。
結論:
- 現在のAI言語モデルは、解剖学試験問題に答える上で顕著な能力を持っており、人間の学生の合格率を超える可能性があります。
- AIのパフォーマンスは急速に向上しているため、教育者はAIの能力を考慮して評価戦略を適応させる必要があります。
- 医学教育における学業の完全性を維持するためには、AIに耐性のある評価方法の開発が不可欠です。
さらに関連する動画
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
3.2K
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
関連する概念動画
Classification of Bones
9.5K
The bones of the human skeletal system are of varied shapes, sizes, and functions. They can be classified based on their shape and function into four major classes: long bones, short bones, flat bones, and irregular bones. Some classifications include a fifth type, the sesamoid bones, as a separate class, whereas others categorize them under short bones.
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
9.5K
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
