Related Experiment Video
Updated: Aug 25, 2025

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
2.6K
A Multigranularity Text Driven Named Entity Recognition CGAN Model for Traditional Chinese Medicine Literatures
Yuekun Ma1,2,3, Yun Liu4, Dezheng Zhang1,3
1School of Computer and Communication Engineering, University of Science and Technology Beijing, Beijing 100083, China.
Computational Intelligence and Neuroscience
|October 17, 2022
Summary
This study introduces a novel Conditional Generation Adversarial Network (MT-CGAN) for Traditional Chinese Medicine Named Entity Recognition (TCM NER). The MT-CGAN model effectively extracts TCM entities from limited annotated data, outperforming existing methods.
Area of Science:
- Computational linguistics
- Bioinformatics
- Traditional Chinese Medicine (TCM) knowledge extraction
Background:
- Extracting Traditional Chinese Medicine (TCM) knowledge from unstructured literature is crucial but challenging.
- Limited large-scale annotated data hinders conventional deep learning models for TCM Named Entity Recognition (NER).
- Existing unsupervised methods often require auxiliary data like domain dictionaries.
Purpose of the Study:
- To develop an effective Traditional Chinese Medicine Named Entity Recognition (TCM NER) model using a small-scale annotated corpus.
- To address the limitations of data scarcity in TCM text knowledge extraction.
- To improve the precision and recall of entity recognition in diverse TCM literature.
Main Methods:
- Proposed a multigranularity text-driven Named Entity Recognition (NER) model named MT-CGAN (Conditional Generation Adversarial Network).
- Designed a multigranularity text features encoder (MTFE) to capture rich semantic and grammatical information.
- Incorporated conditional constraints in the generator and discriminator and introduced seeds from different TCM text types to enhance NER precision.
Main Results:
- The MT-CGAN model demonstrated superior performance compared to baseline methods on four gold-standard datasets.
- Achieved significant improvements in precision (0.24–8.97%), recall (0.89–12.74%), and F1 score (0.01–10.84%).
- The model effectively extracts entities from various TCM literature types, especially those with more entity types and sparsity.
Conclusions:
- The proposed MT-CGAN approach offers a significant advantage for TCM NER, particularly with small-scale and diverse datasets.
- The model's ability to handle texts with higher sparsity and less regular features makes it robust.
- MT-CGAN provides an effective solution for structured TCM knowledge extraction from unstructured texts.

