Related Experiment Video
Updated: Jun 25, 2026

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
3.1K
Large Language Model Evaluation in Traditional Chinese Medicine for Stroke: Quantitative Benchmarking Study
Hulin Long1, Yang Deng2, Yaoguang Guo1
1Hospital of Chengdu University of Traditional Chinese Medicine, Chengdu, Sichuan Province, China.
JMIR Formative Research
|December 11, 2025
Summary
Chinese-centric large language models (LLMs) excel in Traditional Chinese Medicine (TCM) knowledge recall, while general-purpose LLMs perform better in complex reasoning tasks. This study evaluated LLMs using the TCM-SED benchmark.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Traditional Chinese Medicine Research
Background:
- Large language models (LLMs) are increasingly applied in medicine, but evaluating their efficacy in specialized fields like Traditional Chinese Medicine (TCM) presents challenges due to TCM's unique theoretical framework.
- Assessing LLM performance in TCM requires domain-specific benchmarks that capture its distinct knowledge and reasoning requirements.
Purpose of the Study:
- To empirically evaluate the capabilities of different types of large language models (LLMs) within the specialized domain of Traditional Chinese Medicine (TCM) stroke.
- To establish a benchmark dataset for assessing LLM performance in TCM.
Main Methods:
- Developed the Traditional Chinese Medicine-Stroke Evaluation Dataset (TCM-SED), a 203-question benchmark covering diagnosis, treatment, herbal formulas, acupuncture, classical text interpretation, and patient communication.
- Included short-answer, multiple-choice, and essay question formats to assess diverse cognitive levels.
- Evaluated two representative LLMs: GPT-4o (general-purpose) and DeepSeek-R1 (Chinese-centric), using expert-validated gold standard answers.
Main Results:
- DeepSeek-R1 significantly outperformed GPT-4o in objective knowledge recall tasks, achieving over 17% higher accuracy in multiple-choice questions.
- GPT-4o demonstrated superior performance in tasks requiring knowledge integration and complex reasoning, such as interpreting classical TCM texts, scoring 90.5% compared to DeepSeek-R1's 73.5%.
Conclusions:
- Chinese-centric LLMs show advantages in static knowledge tasks within TCM, while general-purpose LLMs excel in dynamic reasoning and content generation.
- The TCM-SED is an effective quantitative tool for evaluating and selecting LLMs for TCM applications, providing a foundation for future model development and optimization.
Keywords:
artificial intelligenceevaluation datasetlarge language modelsmodel evaluationstroketraditional Chinese medicine
