Related Experiment Video
Updated: Oct 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of large language models in the Chinese National Medical Licensing Examinations and Beyond: Scoping
Hui Zong1,2, Jiao Wang1, Jiaxue Cha3
1Joint Laboratory of Artificial Intelligence for Critical Care Medicine, Department of Critical Care Medicine and Institutes for Systems Genetics, Frontiers Science Center for Disease-related Molecular Network, West China Hospital, Sichuan University, Chengdu, China.
Abstract:
Large language models (LLMs) such as GPT-4 and DeepSeek have shown potential in medical education and assessment. In China, the Chinese National Medical Licensing Examination (CNMLE) has become a critical benchmark for evaluating LLMs' professional knowledge and reasoning ability in non-English contexts. This scoping review aims to synthesize research evaluating LLM performance on the CNMLE and related Chinese medical examinations, identifying performance trends, methodological gaps, and future directions. A scoping review was conducted in accordance with PRISMA-ScR guidelines. PubMed and Web of Science were searched on June 12, 2025, using terms related to LLMs and Chinese medical examinations. Studies were included if they evaluated any LLM on national-level Chinese medical exams and reported performance metrics. Two reviewers screened and extracted data. Study quality was assessed using a 12-item checklist covering dataset characteristics, LLM setup, and evaluation methods. Fisher's exact test was used to assess differences in the passing rates of different LLMs. 14 studies were included, covering 51 evaluation records across 8 types of Chinese medical examinations, including CNMLE and specialty exams such as critical care, radiation oncology, and ultrasound medicine. Exam years ranged from 2017 to 2024, with a shift toward using recent exams. A total of 9 LLMs were evaluated, including GPT-3.5, GPT-4, GPT-4o, ERNIE, DeepSeek-R1, Qwen-72B, Baichuan2-7B, Baichuan2-13B and DISC-MedLLM. GPT-4o and DeepSeek-R1 achieved the highest scores of 552 (92.00%) and 523 (87.20%) in the 2024 CNMLE, respectively, both surpassing the passing threshold. In an exploratory comparison, GPT-4 achieved a higher pass rate than GPT-3.5 with a statistically significant difference, although inconsistencies in datasets, model transparency, and evaluation criteria persist across the included studies. LLMs show improving performance on Chinese medical exams, especially newer models. Nevertheless, future research should prioritize standardized benchmarks, multimodal capabilities, and transparent evaluation to ensure meaningful clinical relevance and educational value.