Related Experiment Video
Updated: Jan 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Model Cost and Performance: A Comprehensive Analysis in the Context of the Japan Radiology Board
Takeshi Nakaura1, Naoki Kobayashi1, Kaori Shiraishi1
1Department of Diagnostic Radiology, Graduate School of Medical Sciences.
Closed-source large language models (LLMs) perform better on the Japan Radiology Board Examination (JRBE) than open-source models. Translating questions to English improves LLM accuracy, with higher-cost models showing superior performance.
Area of Science:
- Artificial Intelligence in Medical Education
- Natural Language Processing in Radiology
- Machine Learning for Board Examinations
Background:
- The Japan Radiology Board Examination (JRBE) assesses essential knowledge for radiologists.
- Evaluating the efficacy of Large Language Models (LLMs) in specialized medical examinations is crucial.
- Understanding LLM performance variations based on model type (open-source vs. closed-source) and language is key.
Purpose of the Study:
- To assess the effectiveness of various LLMs in answering JRBE questions.
- To compare the performance of open-source and closed-source LLMs on the JRBE.
- To investigate the impact of language (Japanese vs. English translation) on LLM performance in this context.
Main Methods:
- Administered 315 JRBE questions (2021-2023) to 14 LLMs (7 open-source, 7 closed-source).
- Evaluated models using original Japanese questions and LLM-translated English versions.
- Analyzed performance metrics including median scores, IQR, P-values, and correlation coefficients using Python.
Main Results:
- Closed-source LLMs achieved a 50.2% higher correct response rate than open-source models (44.28% vs. 29.52%, P < 0.001).
- Translating questions to English improved median scores by 21.1% (33.02% to 40.00%, P = 0.005).
- Higher LLM cost correlated positively with accuracy (r = 0.613-0.623, P < 0.020); release date did not significantly impact performance.
Conclusions:
- High-end, closed-source LLMs demonstrate superior performance on the JRBE.
- English translation of questions by LLMs enhances accuracy for these models.
- LLM cost is a significant predictor of performance, while model age is not.
More Related Videos
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
06:16Signal Acquisition, Score Interpretation, and Economics of a Non-Invasive Point-of-Care Test for Coronary Artery Disease
Published on: August 9, 2024
Related Concept Videos
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Radiological Investigation III: Pulmonary Angiogram and PET Scan
Pulmonary Angiogram
A Pulmonary Angiogram is an invasive procedure involving injecting a contrast medium through a catheter threaded into the pulmonary artery or the right side of the heart to visualize the pulmonary vasculature. Computed Tomography (CT) scans have mainly replaced this...
Radiological Investigation I: X-ray and CT
Imaging Studies for Cardiovascular System III: X-Ray
Definition and Purpose
An X-ray, or radiograph, is a non-invasive method that uses ionizing radiation to take images of internal structures. It is mainly used in cardiac imaging to examine the heart, lungs, and major blood vessels, aiming to identify abnormalities in the heart's size, shape, and position, such as heart failure, congenital defects, and vascular...
Imaging Studies II: Positron Emission Tomography and Scintigraphy
Fundamental Principles of PET
Imaging Studies III: Computed Tomography