Related Experiment Video
Updated: Dec 17, 2025

07:35
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
2.0K
Biomedical named entity recognition and linking datasets: survey and our recent development.
Ming-Siang Huang1, Po-Ting Lai2, Pei-Yen Lin3
1Bioinformatics Program, Taiwan International Graduate Program, Institute of Information Science, Academia Sinica, Taipei, Taiwan.
Briefings in Bioinformatics
|July 1, 2020
Summary
This study addresses outdated annotations in biomedical datasets for natural language processing (NLP). Researchers introduce a revised JNLPBA dataset and an Ensembled Biomedical Entity Dataset (EBED) for improved gene, disease, and chemical entity recognition.
Area of Science:
- Bioinformatics
- Computational Biology
- Natural Language Processing
Background:
- Natural language processing (NLP) is crucial for extracting information from biomedical literature.
- Existing biomedical named entity recognition (BNER) datasets face challenges with outdated annotations, inconsistency, and low portability.
- The evolution of NLP techniques necessitates updated and robust benchmark datasets.
Purpose of the Study:
- To review common BNER datasets and identify annotation problems.
- To introduce a revised JNLPBA dataset addressing issues of inconsistency and portability.
- To develop an Ensembled Biomedical Entity Dataset (EBED) for multi-task BNER.
Main Methods:
- Reviewed existing BNER datasets for annotation inconsistencies and portability issues.
- Revised the JNLPBA dataset to resolve identified problems.
- Evaluated the revised dataset's portability using state-of-the-art BNER systems across diverse biomedical texts.
- Extended the revised JNLPBA with additional data sources to create the EBED, incorporating gene, disease, and chemical entity annotations.
Main Results:
- Identified significant annotation problems in commonly used BNER datasets.
- The revised JNLPBA dataset demonstrated improved portability across different biomedical literature types.
- The EBED provides a comprehensive multi-task dataset with 85,000 entity mentions, 25,000 with database identifiers, and 5,000 attribute tags.
Conclusions:
- The developed datasets offer improved resources for training and evaluating BNER systems.
- The EBED facilitates multi-task learning for enhanced biomedical entity recognition.
- These datasets contribute to more robust and reliable information retrieval from biomedical publications.

