Related Experiment Video
Updated: Aug 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A Rasch-based analysis comparing the performance of updated multimodal large language models and surgeons on the
Yuji Miyamoto1,2, Takeshi Nakaura3, Ayane Kawata4
1Department of Gastroenterological Surgery, Graduate School of Medical Sciences, Kumamoto University, 1-1-1 Honjo, Kumamoto, 860-8556, Japan. miyamotoyuji@kumamoto-u.ac.jp.
Purpose:
To compare four multimodal large language models (LLMs) with surgeons on the 2023 Japanese Surgical Specialist Examination using item-level surgeon correct answer rates as the benchmark.
Methods:
In this retrospective cross-sectional study, GPT-4.1, Claude Opus 4, Gemini 2.5 Pro, and o3 Pro were evaluated using 98 valid multiple-choice items, including 43 image-based and 55 text-only questions. The accuracy was examined overall, by image presence, and by subspecialty. Physician-anchored Rasch modeling placed LLMs and surgeons on a common latent scale, and logistic regression assessed the association between surgeon accuracy and LLM correctness at the item-level.
Results:
The accuracy ranged from 77.6% for GPT-4.1 to 85.7% for o3 Pro. All models showed lower accuracy on image-based items than on text-only items. A Rasch analysis showed that all LLMs remained below the surgeons' overall, with relative abilities ranging from - 1.29 to - 0.74. Performance varied according to difficulty and subspecialty. Gastroenterology was consistently the weakest domain, whereas some models matched or exceeded the surgeon benchmark in selected areas, including respiratory, pediatrics, breast/endocrine, and emergency/anesthesiology.
Conclusions:
For this previously analyzed item set, updated multimodal LLMs achieved an overall accuracy below the surgeon benchmark, with the largest deficits on image-based items, supporting supervised and domain-specific use.
