Related Experiment Video
Updated: Jan 28, 2026

Enema of Traditional Chinese Medicine for Patients with Severe Acute Pancreatitis
Published on: January 27, 2023
Supporting postgraduate exam preparation with large language models: implications for traditional Chinese medicine
Baifeng Wang1, Meiwei Zhang2, Zhe Wang3
1Dongzhimen Hospital, Beijing University of Chinese Medicine, Beijing, China.
Introduction:
In China, the medical education system features multiple co-existing levels, with higher education often leading to better job prospects. In career advancement-especially for entry into competitive urban hospitals-the postgraduate examination often plays a more decisive role than the licensing examination. The application of Large Language Models (LLMs) in Traditional Chinese Medicine (TCM) has rapidly expanded. TCM theories possess distinct scientific features, requiring LLMs to demonstrate advanced information processing and comprehension abilities in a Chinese context. While LLMs have shown strong performance in many countries' licensing examinations, their performance in selective TCM examinations remains underexplored. This study aimed to evaluate and compare the performance of Ernie Bot, ChatGLM, SparkDesk, and GPT-4 on the 2023 Chinese Postgraduate Examination for TCM (CPE-TCM), and explore their potential in supporting TCM education and academic development.
Methods:
We assessed the performance of four LLMs using the 2023 CPE-TCM as a test set. Exam scores were calculated to evaluate subject-specific performance. Additionally, responses were qualitatively analyzed based on logical reasoning and the use of internal and external information.
Results:
Ernie Bot and ChatGLM achieved accuracy rates of 50.30 and 46.67%, respectively, both above the passing score. Statistically significant differences in subject-specific performance were observed, with the highest scores in the medical humanistic spirit module. ChatGLM and GPT-4 provided logical explanations for all responses, while Ernie Bot and SparkDesk showed logical reasoning in 98.2 and 43.6% of responses, respectively. ChatGLM and GPT-4 incorporated internal information in all explanations, whereas SparkDesk rarely did. Over 60% of responses from Ernie Bot, ChatGLM, and GPT-4 included external information, which did not significantly differ between correct and incorrect answers. In SparkDesk, the presence of internal or external information was significantly associated with answer correctness (P < 0.001).
Discussion:
Ernie Bot and ChatGLM surpassed the passing threshold for postgraduate selection, reflecting solid TCM expertise. LLMs demonstrated strong capabilities in logical reasoning and integration of background knowledge, highlighting their promising role in enhancing TCM education.
Related Concept Videos
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition
Self-Help Support Groups
Accessibility and Cost-Effectiveness
One of the primary strengths of self-help...
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...

