推理优化的大型语言模型在板式骨科考试中达到近专家准确度:对702个多选题进行多模型比较
Pedro Diniz1,2,3, Takuji Yokoe4, Felix C Öttl5
1Department of Orthopaedic Surgery, Centre Hospitalier Universitaire Brugmann, Brussels, Belgium.
概括
较新的推理优化的大型语言模型 (LLM) 在骨科考试问题上显著优于GPT-4. 这些先进的LLM显示出更高的精度和更好的校准,尽管延迟和成本仍然是临床使用的考虑因素.
科学领域:
- 人工智能在医学中的应用
- 医疗教育 技术 技术 医学教育
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 越来越多地被用于医学应用.
- 评估LLM在特殊医学知识 (如骨科) 上的表现至关重要.
- 将高级推理LLM与像GPT-4这样的既定模型进行比较是了解进步的必要条件.
研究的目的:
- 为了比较七个LLM的精度,校准,可重复性和运营成本,对骨科多项选择题 (MCQ) 进行了比较.
- 评估具有高级推理能力的新型LLM的性能.
- 量化这些模型与GPT-4相比的性能增长.
主要方法:
- 702个独特的骨科MCQ从Orthobullets中提取出来.
- 这些问题由OpenAI o3,Anthropic Claude Sonnet 4,Claude Opus 4,Google Gemini 2.5 Pro,GPT-4,GPT-4o和Gemma 3 27B进行了分析.
- 主要结果是整体准确性;次要结果包括校准,可重复性,延迟和成本.
主要成果:
- 推理优化的LLM得分至少比GPT-4高14个百分点 (OpenAI o3的准确率为93.6%).
- 准确性随着问题难度的下降而下降,但推理优势在所有层面上都存在.
- 与GPT-4 (ECE=0.215) 相比,Claude Opus 4显示出更优异的校准 (ECE=0.023).与GPT-4相比,Claude Opus 4显示出更优异的校准 (ECE=0.023).
结论:
- 推理优化的LLM在骨科考试问题上实现了高精度和改进的信心校准.
- 尽管取得了进展,但模型随机性和显著的延迟成本权衡等问题可能会阻碍广泛的临床采用.
- 这些发现凸显了人工智能在医学教育和评估方面的快速进展.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.2K
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
3.1K
相关概念视频
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
261
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
261
Multiple Comparison Tests
4.4K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
4.4K
Classification of Bones
9.4K
The bones of the human skeletal system are of varied shapes, sizes, and functions. They can be classified based on their shape and function into four major classes: long bones, short bones, flat bones, and irregular bones. Some classifications include a fifth type, the sesamoid bones, as a separate class, whereas others categorize them under short bones.
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
9.4K
