Related Experiment Video
Updated: Aug 12, 2026

Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
Published on: September 5, 2019
[Comparative study on the application of large language models in pre-exam learning for the Chinese dental licen-sing
Xiao Pang1, Chang Liu1, Jiahao Fan2
1State Key Laboratory of Oral Diseases & National Center for Stomatology & National Clinical Research Center for Oral Diseases & Dept. of Information Management, West China Hospital of Stomatology, Sichuan University, Chengdu 610041, China.
Objectives:
This study aims to compare the performance of different large language models (LLMs) in pre-exam learning for the Chinese dental licensing examination, with a focus on evaluating their differences in answering, explanation, and teaching effectiveness, to provide a reference for the application of LLMs in dental education.
Methods:
Three evaluation scenarios were designed: selecting correct answers, providing answer explanations, and adversarial testing. DeepSeek-R1, Qwen 2.5-MAX, Doubao 1.5 Pro, Xinghuo Spark-X1, ERNIE 4.0 Turbo, GPT-4o, and Huaxi Zhilian were selected for comparative testing. Evaluation metrics included accuracy, net accuracy, and pedagogical effectiveness.
Results:
In the scenario of selecting correct answers, all LLMs exceeded the passing threshold, with Huaxi Zhilian achieving the highest accuracy (84%). In the answer explanation scenario, Huaxi Zhilian demonstra-ted the highest net accuracy (92%), followed by DeepSeek-R1 (89%), among models. Regarding pedagogical effectiveness, Huaxi Zhilian ranked highest in relevance, practicality, and clarity, whereas GPT-4o led in conciseness, among the investigated LLMs. In adversarial testing, Huaxi Zhilian and DeepSeek-R1 exhibited the smallest declines in accuracy and net accuracy, respectively, among the tested models.
Conclusions:
In pre-exam learning for the Chinese dental licensing examination, knowledge-enhanced LLMs specifically optimized for dentistry (e.g., Huaxi Zhilian) outperform reasoning LLMs pretrained on general Chinese corpora (e.g., DeepSeek-R1) and those primarily trained on English corpora (e.g., GPT-4o). However, the performance of all models declines under adversarial conditions. Future research should focus on addressing identified weaknesses to enhance the utility of LLMs in dental education further.
