Related Experiment Video
Updated: Sep 30, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comprehensive Evaluation of Large Language Models on Four Core Medical School Courses: A Cross-Sectional Comparative
Kai Zhang1, Wenjing Yang2, Wei Zheng2
1College of Artificial Intelligence Medicine, Chongqing Medical University, Chongqing, People's Republic of China.
Purpose:
Large language models (LLMs) have demonstrated remarkable potential in medical education, yet their performance on discipline-specific medical school course examinations remains incompletely characterized. This study evaluated six contemporary LLMs on final examinations for four core medical school courses: Surgery, Musculoskeletal System Diseases, Digestive System Diseases, and Respiratory System Diseases.
Methods:
A total of 400 multiple-choice questions (100 per course) were administered to each model. Responses were scored against official answer keys and compared to the performance of 312 medical students. Each question was tested three times per model to assess response consistency and reproducibility.
Results:
All six LLMs achieved mean scores exceeding 93% across all four courses, substantially outperforming the student cohort mean of 71.8%, with all models scoring above the 99.5th percentile of the student score distribution. ChatGPT 5.5 achieved the highest overall accuracy (97.0%), followed by Claude Opus 4.7 (96.3%) and DeepSeek V4 (95.8%). Qwen 3.7 (95.3%) and GLM 5.2 (94.3%) demonstrated comparable performance, while Doubao 2.1 (93.8%) showed the lowest but still exceptional accuracy. A sensitivity analysis revealed that accuracy on the 12 visual-containing questions was substantially lower (50.0%) than on text-only questions (96.8%). All models demonstrated high response consistency (κ = 0.82-0.95) and reproducibility (>95%).
Conclusions:
Contemporary LLMs can achieve near-perfect performance on medical school course examinations, substantially exceeding average student performance. While this capability suggests significant potential for LLMs as supplementary educational tools, it also raises important concerns regarding assessment integrity and the appropriate role of AI in medical training.
