Benchmark-Based Evaluation of ChatGPT and Gemini in Radiation Oncology: Performance, Limitations, and Challenges for
Yousef Mohammed1,2, Ronny Kruschel3, Nizar Alshammas4
1Radiation Oncology and Radiotherapy, SRH Zentralklinikum Suhl, Suhl, DEU.
Abstract:
Background Large language models (LLMs) are increasingly being explored for medical applications, including clinical decision support and oncology education. However, their performance in radiation oncology remains insufficiently characterized. Methods This study evaluated and compared the performance of ChatGPT (GPT-5.2; OpenAI, San Francisco, USA) and Gemini (Ultra/Pro; Google, Mountain View, USA) in radiation oncology. A benchmark consisting of 70 multiple-choice questions covering clinical oncology, radiation physics, and radiobiology was used to assess general knowledge. In addition, 25 clinically relevant open-ended questions were independently evaluated by three radiation oncologists using five-point Likert scales for correctness and usefulness. A mixed-effects model was applied to analyze performance. Results ChatGPT achieved an accuracy of 94.3%, while Gemini achieved 97.1% in the multiple-choice assessment. For open-ended questions, both models received similarly high ratings, with mean correctness scores of 4.71 and 4.67 and mean usefulness scores of 4.63 and 4.64 for ChatGPT and Gemini, respectively. Mixed-effects analysis demonstrated a significant effect of question type on both correctness and usefulness, whereas no significant differences between models were observed. Although most responses were rated as good or very good, limitations became apparent in more complex clinical scenarios requiring prioritization and individualized decision-making. Minor discrepancies between benchmark answers and current clinical evidence were also identified. Conclusions Both ChatGPT and Gemini demonstrated high performance on benchmark-based radiation oncology assessments. While the generated responses were generally accurate, limitations remained in complex clinical scenarios requiring nuanced clinical judgment. Further studies are needed to determine whether such performance translates into meaningful clinical utility in real-world radiation oncology practice.


