Related Experiment Video
Updated: May 10, 2026

Gene Regulation and Targeted Therapy in Gastric Cancer Peritoneal Metastasis: Radiological Findings from Dual Energy CT and PET/CT
Published on: January 22, 2018
A Comparative Analysis of GPT-4o and ERNIE Bot in a Chinese Radiation Oncology Exam
Weiping Wang1, Jingxuan Fu2, Yiming Zhang3
1Department of Radiation Oncology, Peking Union Medical Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College, No. 1 Shuaifuyuan Wangfujing, Beijing, 100730, China.
Abstract:
Large language models (LLMs) are increasingly utilized in medical education and practice, yet their application in niche fields such as radiation oncology remains underexplored. This study evaluates and compares the performance of OpenAI's GPT-4o and Baidu's ERNIE Bot in a Chinese-language radiation oncology examination. We employed the Chinese National Health Professional Technical Qualification Examination (Intermediate Level) for Radiation Oncology, using a question bank of 1128 items across four sections: Basic Knowledge, Relevant Knowledge, Specialized Knowledge, and Practice Competence. A passing score required an accuracy rate of 60% or higher in all sections. The models' responses were assessed for accuracy against standard answers, with key metrics including overall accuracy, section-specific performance, case analysis performance, and accuracy consensus between the models. The overall accuracy rates were 79.3% for GPT-4o and 76.9% for ERNIE Bot (p = 0.154). Across the four sections, GPT-4o achieved accuracy rates of 82.1%, 84.6%, 78.6%, and 60.9%, respectively, while ERNIE Bot achieved 81.6%, 73.9%, 77.9%, and 69.0%. In the Relevant Knowledge section, GPT-4o achieved significantly higher accuracy (p = 0.002), while no significant differences were found in the other three sections. Across various question types-including single-choice, multiple-answer, case analysis, non-case analysis, and different content areas of case analysis-both models exhibited satisfied accuracy, and ERNIE Bot achieved accuracy rates that were comparable to GPT-4o. The accuracy consensus between the two models was 84.5%, significantly exceeding the individual accuracy rates of GPT-4o (p = 0.003) and ERNIE Bot (p < 0.001). Both GPT-4o and ERNIE Bot successfully passed the highly specialized Chinese-language medical examination in radiation oncology and demonstrated comparable performance. This study provides valuable insights into the application of LLMs in Chinese medical education. These findings support the integration of LLMs in medical education and training within specialized, non-English-speaking contexts.
Related Concept Videos
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...
Imaging Studies II: Positron Emission Tomography and Scintigraphy
Fundamental Principles of PET
Radiological Investigation III: Pulmonary Angiogram and PET Scan
Pulmonary Angiogram
A Pulmonary Angiogram is an invasive procedure involving injecting a contrast medium through a catheter threaded into the pulmonary artery or the right side of the heart to visualize the pulmonary vasculature. Computed Tomography (CT) scans have mainly replaced this...
