Related Experiment Video
Updated: Oct 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Reasoning and Comparator Large Language Models on Nephrology Multiple-choice Questions
Fumiya Kitano1, Mamoru Masaki1, Daisuke Ichikawa1
1Division of Nephrology and Hypertension, Department of Internal Medicine, St. Marianna University School of Medicine, Kawasaki, Japan.
Introduction:
Performance of large language models in medicine is improving; yet, it remains unclear how the advantage of reasoning models depends on task characteristics in nephrology.
Methods:
We evaluated four large language models in two families-OpenAI (GPT-5 reasoning and GPT-4o comparator) and Google (Gemini 2.5 Pro reasoning and Gemini 2.0 Flash comparator)-on 209 self-assessment questions for nephrology board renewal published by the Japanese Society of Nephrology. Questions were categorized by question type (general vs. clinical), taxonomy (recall, interpretation, and problem-solving), and image inclusion (non-image vs. image). Models were assessed via application programming interface; no generation parameters were configured; only the model was selected. Images were provided as PNG files. Accuracy used Wilson 95% confidence intervals (CIs); paired comparisons used McNemar's exact test. Primary analyses used logistic generalized linear mixed models with fixed effects, random intercepts, and prespecified interactions.
Results:
Overall accuracy was 87.6% (183/209, 95% CI 82.4-91.4) for GPT-5 and 83.7% (175/209, 95% CI 78.1-88.1) for Gemini 2.5 Pro versus 69.9% (146/209, 95% CI 63.3-75.7) for GPT-4o and 62.7% (131/209, 95% CI 55.9-69.0) for Gemini 2.0 Flash. Paired analyses favored the reasoning models, with odds ratios of 6.29 for OpenAI and 7.29 for Google (both p <0.001). Adjusted odds ratios for reasoning versus comparator models at the reference strata (general, recall, and non-image questions) were 5.00 for OpenAI and 7.28 for Google (both p <0.001). Interactions showed stronger effects in clinical questions for OpenAI and taxonomy-dependent effects for Google; no significant modification by image inclusion was observed.
Conclusions:
The reasoning models tested here outperformed their comparator models on nephrology board-style multiple-choice questions, with advantages that varied by question type. These findings support further evaluation of reasoning models as supplementary tools for nephrology board-recertification preparation, while underscoring that direct clinical use requires task-specific validation.