Related Experiment Video
Updated: Aug 10, 2026

Operational and Intervention Effects of Targeted Tuina in Lumbar Intervertebral Disc Degeneration Model Rabbits
Published on: July 21, 2023
Performance Comparison Between Domestic and International Large Language Models in Patient Education for Chinese
Quan Zhang1,2, Ruizhong Yan2, Xiaoliang Liu1
1Department of Orthopedics, The Second Clinical Medical College of Shanxi Medical University, The second hospital of Shanxi Medical university, Taiyuan, China.
Objective:
This study aims to systematically evaluate and compare the performance of different large language models (LLMs) in providing medical advice for lumbar disc herniation (LDH) within the Chinese clinical context, addressing the current lack of standardized assessment tools for orthopedic guidance in China.
Methods:
We constructed a standardized dataset covering diagnosis, treatment, surgical risks, and healthcare navigation. Using a unified prompt, we tested eight models: four international (ChatGPT-4o, Gemini-3.0-Pro, Claude-4.0, Grok-4-auto) and four Chinese (DeepSeek, Doubao, Qwen-Max, Kimi). Senior spine surgeons evaluated responses in a blinded manner using the Mika scale (accuracy), modified DISCERN (reliability), EQIP, and GQS. Readability was assessed via the Ludong University Text Grading Platform.
Results:
Significant performance differences emerged among models (P < 0.001). DeepSeek and Doubao achieved the highest median accuracy (2.00; IQR 2.00-2.00). We observed a distinct "competency divergence": international models displayed stronger safety guardrails and logical robustness in complex reasoning, whereas Chinese models demonstrated unique advantages in service contextualization and local resource navigation. Currently, no single model perfectly integrates global medical evidence with local delivery contexts.
Conclusion:
DeepSeek, Doubao, and Gemini are the preferred LLMs for LDH patient education in China. While leading domestic models challenge the stereotype of international superiority in medical logic, the general scarcity of "excellent" responses (<25%) indicates LLMs remain auxiliary tools. Future optimization should prioritize Retrieval-Augmented Generation (RAG) anchored in local clinical guidelines to ensure safe deployment.
Related Concept Videos
Herniated Intervertebral Disc l: Introduction
Degenerative Disc Disease I: Introduction