Related Experiment Video
Updated: Jul 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance stability despite iteration: evaluating DeepSeek and ChatGPT on Chinese medical licensing examinations
Zhiheng Wang1, Yifan Qin1, Jin Wu1
1Department of Anesthesiology, Affiliated Hospital of Jiangsu University, Zhenjiang, China.
Background:
Large Language Models (LLMs) hold substantial potential in medical education. In our previous work, we evaluated the performance of DeepSeek-R1 and ChatGPT-4o on the Chinese National Medical Licensing Examination (CNMLE). Following performance upgrades in December 2025, DeepSeek-V3.2 and ChatGPT-5.2 were released. This study aimed to longitudinally assess the performance evolution of DeepSeek (R1 vs. V3.2) and ChatGPT (4o vs. 5.2) using the 2024 CNMLE as a baseline and to explore their performance on the 2025 CNMLE.
Methods:
We tested DeepSeek-V3.2 and ChatGPT-5.2 on 600 multiple-choice questions from the written part of the 2024 CNMLE, and compared the results with historical data from DeepSeek-R1 and ChatGPT-4o. The questions consisted of 4 units and were divided into low-difficulty and high-difficulty groups according to different difficulty levels. Additionally, the two latest LLMs were assessed on 600 questions from the written part of the 2025 CNMLE.
Results:
In the 2024 CNMLE, overall accuracy for the DeepSeek series (R1 vs. V3.2: 92.0% vs. 91.0%) and ChatGPT series (4o vs. 5.2: 87.2% vs. 89.3%) showed no significant differences (all p > 0.05). The accuracy gap between the two series narrowed from 4.8% (DeepSeek-R1 vs. ChatGPT-4o, p < 0.05) to 1.7% (DeepSeek-V3.2 vs. ChatGPT-5.2, p = 0.332). Subgroup analyses by unit and difficulty revealed no statistically significant differences, either in longitudinal comparisons between successive versions within each series or in cross-sectional comparisons between the latest versions of the two series. On the 2025 CNMLE, DeepSeek-V3.2 achieved significantly higher overall accuracy than ChatGPT-5.2 (94.7% vs. 89.3%, p < 0.05), with superior performance in Unit 2 and in both the low- and high-difficulty groups (p < 0.05). Error analysis showed no significant differences across model versions in error type classification or clinical risk rating, despite a substantial proportion of high-risk clinical errors.
Conclusion:
Benchmarked against the 2024 CNMLE, iterative updates did not yield significant performance gains for either LLM series. However, DeepSeek-V3.2 demonstrated a performance advantage on the 2025 CNMLE.