放射線治療における大規模言語モデルの能力評価:日本の専門認定試験を通じて
Noriyuki Kadoya1, Yoshiyuki Takahashi1, Seiya Koga1
1Department of Radiation Oncology, Tohoku University Graduate School of Medicine, 1-1 Seiryo-machi, Aoba-ku, Sendai, Miyagi 980-8574, Japan.
Abstract:
Large language models (LLMs), such as ChatGPT and Grok, have rapidly advanced in natural language understanding and are increasingly being applied to specialized fields, including medicine. In this study, we evaluated the domain-specific knowledge of LLMs in radiotherapy by assessing their performance on three certification examinations in Japan: the Japanese Medical Physicist Examination, the Japanese Board Examination for Radiologists and the Japanese Board Examination for Radiation Oncologists. We assessed five LLMs-ChatGPT-5, ChatGPT-5 Pro, Grok 4, Grok 4 heavy and Gemini 2.5 Pro-by inputting all multiple-choice questions from these exams into each model and recording their responses. The AI-generated answers were compared with reference answers determined by experienced medical physicists and radiation oncologists. The results demonstrated average accuracies of 84.7 ± 2.0% (ChatGPT-5), 94.7 ± 2.1% (ChatGPT-5 Pro), 78.4 ± 1.2% (Grok 4), 81.6 ± 2.2% (Grok 4 heavy) and 88.9 ± 1.2% (Gemini 2.5 Pro). All models achieved over 75% accuracy, with ChatGPT-5 Pro consistently outperforming others, attaining an average accuracy exceeding 90% across all examinations. These findings highlight the strong potential of advanced LLMs, particularly ChatGPT-5 Pro, for future integration into radiotherapy-related applications such as automated contouring and treatment planning support.
さらに関連する動画
関連する概念動画
Radiation: Applications
The average...
Biological Effects of Radiation
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Cancer Survival Analysis


