Related Experiment Video
Updated: Aug 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A comparative study of the performance of different large language models in the Chinese National Pharmacist
Can Huang1, Yanfang Sun1, Wei Liu1
1Beijing Youan Hospital, Capital Medical University, Beijing, China.
Objective:
To systematically evaluate the overall performance, subject-based differences, and question type adaptability of five mainstream large language models (ChatGPT, DeepSeek, Kimi, Qwen, and Doubao) in the Chinese National Pharmacist Licensing Examination (CNPLE), and to explore their feasibility as auxiliary tools for pharmaceutical examinations.
Methods:
A cross-sectional comparative study design was adopted. The practice questions of the 2024 CNPLE were used as the evaluation dataset, covering 480 standardized questions across four subjects: Pharmaceutical Professional Knowledge (I), Pharmaceutical Professional Knowledge (II), Comprehensive Knowledge and Skills of Pharmacy, and Pharmaceutical Administration and Regulation. The question types included Type A (single best choice), Type B (matching choice), Type C (comprehensive analysis), and Type X (multiple choice). Standardized prompts were used for independent tests using the official web versions of each model with default parameters. Each question was input separately in a new conversation session to avoid contextual interference. Taking the official standard answers as the gold standard, the subject accuracy rate, question type accuracy rate, and overall accuracy rate of each model were calculated. To compare the overall performance among the five models, Cochran's Q test was applied. Post-hoc pairwise comparisons were performed using McNemar's test with Bonferroni correction for multiple comparisons.
Results:
All five models exceeded the 60% passing score threshold of the CNPLE. The overall accuracy ranking was: Kimi (89.58%) > Doubao (88.96%) > DeepSeek (87.29%) > Qwen (77.92%) > ChatGPT (72.50%). Cochran's Q test showed a statistically significant difference in the overall accuracy among the five models (Q = 111.39, df = 4, P < 0.001). Pairwise comparisons showed no significant differences among Kimi, Doubao and DeepSeek (P > 0.05), while all three performed significantly better than Qwen and ChatGPT (P < 0.001). At the subject level, all models achieved the best performance in Pharmaceutical Professional Knowledge (II) (average accuracy 89.50%) and relatively weak performance in Pharmaceutical Administration and Regulation (average accuracy 76.50%). At the question type level, Type C questions yielded the highest average accuracy (90.00%), whereas Type X questions had the lowest average accuracy (69.00%).
Conclusion:
In this single-run evaluation, the Chinese large language models tested achieved higher overall accuracy than ChatGPT under the same conditions. Among them, Kimi, Doubao and DeepSeek have reached an excellent performance level. Different models present differentiated advantages across subjects and question types. Regulatory subjects and Type X (multiple-answer) questions are common challenges for all models. The findings indicate that mainstream LLMs possess considerable potential as auxiliary tools for the CNPLE, and can provide intelligent support for pharmaceutical education and examination preparation.

