Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models for the National Radiological Technologist Licensure Examination in Japan: Cross-Sectional
Toshimune Ito1,2,3, Toru Ishibashi1, Tatsuya Hayashi1,2
1Department of Radiological Technology, Faculty of Medical Technology, Teikyo University, 2-11-1 Kaga, Itabashi-ku, Tokyo, 173-8605, Japan, +81-3-3964-7053.
Large language models (LLMs) can generate accurate multiple-choice questions for radiological technologist exams. While OpenAI o3 excelled in accuracy and content alignment, further refinement is needed for wording clarity and instructional usefulness.
Area of Science:
- Medical Education Technology
- Artificial Intelligence in Healthcare
- Radiological Sciences
Background:
- Mock examinations are crucial for health professional education and licensure preparation.
- Instructor-written multiple-choice questions often face challenges in consistency and clarity.
- Large language models (LLMs) show promise for medical exam item development, but their educational quality requires thorough evaluation.
Purpose of the Study:
- To identify the most accurate large language model (LLM) for generating questions for the Japanese National Examination for Radiological Technologists.
- To utilize the top-performing LLM to create blueprint-aligned multiple-choice questions.
- To assess the educational quality of LLM-generated questions through expert review.
Main Methods:
- Four LLMs (OpenAI o3, o4-mini, o4-mini-high, Gemini 2.5 Flash) were tested on the 77th Japanese National Examination for Radiological Technologists.
- Accuracy was evaluated for all items and a subset of 173 non-image items.
- The best model (o3) generated 192 new items, which were rated by subject-matter experts on difficulty, factual accuracy, content coverage, wording, and instructional usefulness.
Main Results:
- OpenAI o3 demonstrated the highest accuracy (90.0% overall, 92.5% on non-image items), significantly outperforming o4-mini.
- Generated items received high expert ratings for difficulty (4.29), factual accuracy (4.18), and content coverage (4.73).
- Ratings for wording appropriateness (3.92) and instructional usefulness (3.60) were lower, not consistently meeting the 'adoptable' threshold (≥4).
Conclusions:
- OpenAI o3 effectively generates radiological licensure items meeting national standards for difficulty, accuracy, and blueprint alignment.
- Wording clarity and pedagogical specificity of explanations require further editorial refinement for optimal educational quality.
- A hybrid approach, with LLMs drafting items and faculty refining them, offers a practical workflow for scalable, high-quality item bank development.
More Related Videos
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023