Related Experiment Video
Updated: Oct 10, 2026

A Study on an Intelligent Diagnosis and Treatment Assistant System for Acupuncture in Diminished Ovarian Reserve Based on a Knowledge Graph
Published on: May 29, 2026
Nested named entity recognition in Chinese electronic medical records enhanced by part-of-speech and entity boundary
Jianmei Wang1, Yupeng Liao1, Zhaoben Xie1
1School of Medical and Information Engineering, Gannan Medical University, Ganzhou, Jiangxi Province, China.
Abstract:
ObjectiveNested named entity recognition in Chinese electronic medical records (EMRs) is challenging because entity spans overlap and have ambiguous boundaries. We evaluated whether part-of-speech (POS) cues and automatically extracted entity boundary knowledge could improve a RoBERTa-GlobalPointer framework for this task.MethodsThis retrospective study used public, de-identified benchmark datasets. We evaluated a RoBERTa-wwm-ext-large-GlobalPointer framework with two enhancement branches: token-aligned POS embeddings derived from jieba. posseg and an entity boundary knowledge base (EBKB) constructed from entity head, tail and length patterns. The primary experiment used Chinese Medical Entity Extraction version 2 (CMeEE-V2), and supplementary test-set analyses used cEHRNER and CCKS2019.ResultsOn the CMeEE-V2 test set, the best configuration achieved 72.45% precision, 73.02% recall and 72.60% F1, outperforming the comparators. Ablation results indicated that boundary knowledge was the principal source of improvement, whereas POS features had a smaller effect. The cEHRNER test-set analysis provided additional cross-dataset evidence. Across five CCKS2019 test-set runs, F1 was 0.8152 ± 0.0096 for the knowledge-enhanced model and 0.7054 ± 0.0435 for the baseline (paired p=0.0043).ConclusionAutomatically extracted boundary knowledge was the main driver of improvement in the framework, while POS features acted as auxiliary cues. Findings were consistent across primary and supplementary test-set evaluations.