Related Experiment Video
Updated: Jul 25, 2025

13:19
Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
9.2K
EduNER: a Chinese named entity recognition dataset for education research
Xu Li1, Chengkun Wei1, Zhuoren Jiang2
1College of Computer Science and Technology, Zhejiang University, 38 Zheda Rd., Hangzhou, 310027 Zhejiang China.
Neural Computing & Applications
|June 26, 2023
Summary
Researchers developed EduNER, a new Chinese dataset for named entity recognition (NER) in education. This resource aims to advance domain-specific NER models by providing diverse, expert-annotated data.
Area of Science:
- Natural Language Processing
- Information Extraction
- Educational Technology
Background:
- Domain-specific Named Entity Recognition (NER) requires high-quality, targeted datasets.
- Existing datasets often lack coverage for the specialized terminology and context within the education domain.
- The development of robust education-oriented NER models is hindered by the absence of a comprehensive, publicly available dataset.
Purpose of the Study:
- To introduce EduNER, a novel, high-quality, domain-oriented Chinese Named Entity Recognition dataset specifically for the education sector.
- To facilitate the development and evaluation of advanced Chinese NER models tailored to educational contexts.
- To provide a benchmark dataset for future research in education-specific NLP tasks.
Main Methods:
- Data collection from diverse sources including textbooks, academic papers, and education-related web pages spanning 2012-2021.
- Expert-driven schema definition for 16 education-specific entity types.
- Annotation by trained annotators using a collaborative labeling platform, resulting in 11k+ sentences and 35,731 annotated entities.
Main Results:
- The EduNER dataset comprises 16 distinct entity types, over 11,000 sentences, and 35,731 annotated entities.
- Statistical analysis and comparison reveal EduNER's unique characteristics relative to open-domain and other domain-specific NER datasets.
- Evaluation of sixteen state-of-the-art models on EduNER demonstrates its utility for task validation and highlights potential for model improvement.
Conclusions:
- EduNER is presented as the first publicly available dataset for Chinese NER in the education domain.
- The dataset's quality, diversity, and expert annotation are expected to significantly promote the development of education-oriented NER models.
- This resource offers a valuable foundation for advancing NLP applications within the educational technology landscape.

