Related Experiment Videos
Benchmarking large language models on a Chinese radiation oncology technology examination-preparation question set:
Heling Zhu1, Yongguang Liang1, Xinyu Long2
1Department of Radiation Oncology, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China.
Purpose:
This study evaluated GPT-4o, GPT-5.4, and two DeepSeek platform configurations using 1, 053 examination-preparation questions from a publicly and commercially available 2025 exercise collection for the Chinese National Radiation Oncology Technology Qualification Examination (Intermediate Level). We assessed accuracy rates, identical-response coverage, accuracy consensus, and model-reported latency estimates.
Methods:
Four configurations-GPT-4o, GPT-5.4, DS-Fast, and DS-Expert-were tested using an identical zero-shot Chinese prompt. Accuracy rates were summarized with Wilson 95% confidence intervals. Pairwise differences in the primary Chinese-language analysis were assessed using exact McNemar's tests with Holm adjustment. A descriptive language-controlled analysis evaluated English translations of the same questions using an equivalent English prompt.
Results:
Overall accuracy rates were 56.7% (95% CI, 53.7-59.7) for GPT-4o, 66.1% (63.2-68.9) for GPT-5.4, 98.4% (97.4-99.0) for DS-Fast, and 99.8% (99.3-99.9) for DS-Expert. All six overall pairwise comparisons remained significant after Holm adjustment. DS-Fast and DS-Expert produced identical responses for 1, 034 of 1, 053 questions (98.2%), all of which were correct. GPT-5.4 and DS-Expert showed 100% conditional accuracy among shared responses, but their identical-response coverage was only 65.9% (694/1, 053). Model-reported latency estimates were lowest for DS-Fast (0.68 ± 1.00 seconds) and highest for DS-Expert (3.69 ± 2.24 seconds). Under the English-translated condition, overall accuracy rates were 62.6%, 61.5%, 66.5%, and 76.9%, respectively. DS-Expert remained the most accurate configuration, although its advantage over the GPT models was reduced.
Conclusion:
DeepSeek configurations outperformed the GPT models on this Chinese-language examination-preparation question set, with DS-Expert achieving the highest accuracy. However, performance varied substantially by question language. These findings support further evaluation for answer verification and practice-question review, but they do not establish educational effectiveness.