Related Experiment Video
Updated: Nov 18, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
A pre-training and self-training approach for biomedical named entity recognition.
Shang Gao1, Olivera Kotevska2, Alexandre Sorokine3
1Computational Sciences and Engineering Division, Oak Ridge National Laboratory, Oak Ridge, TN, United States of America.
Transfer learning and self-training significantly improve named entity recognition (NER) in biomedical settings with limited labeled data. These methods allow models to achieve high performance using substantially less expert annotation.
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Computational Biology
Background:
- Named Entity Recognition (NER) is crucial for scientific literature mining.
- Current NER models often require extensive labeled data, limiting their use in specialized biomedical domains.
- Acquiring expert annotations for biomedical text is challenging and costly.
Purpose of the Study:
- To investigate the effectiveness of transfer learning and semi-supervised self-training for biomedical NER with limited labeled data.
- To enhance NER model performance in resource-scarce biomedical applications.
- To reduce the dependency on large annotated datasets for effective NER.
Main Methods:
- Pre-training BiLSTM-CRF and BERT models on large general biomedical NER corpora (e.g., MedMentions, Semantic Medline).
- Fine-tuning pre-trained models on specific target NER tasks with limited labeled data (250-2000 samples).
- Applying semi-supervised self-training using unlabeled data to further improve model performance.
Main Results:
- Combining transfer learning and self-training achieved performance comparable to models trained on 3x-8x more labeled data for common biomedical entities (UMLS).
- The approach successfully boosted performance in low-resource scenarios with rare entity types not covered by UMLS.
- Demonstrated significant improvements in NER accuracy despite data scarcity.
Conclusions:
- Transfer learning coupled with self-training offers a powerful strategy for developing effective biomedical NER systems with minimal labeled data.
- This methodology addresses the bottleneck of data annotation in specialized scientific domains.
- The proposed approach is adaptable to both common and rare biomedical entity recognition tasks.
More Related Videos
05:22Author Spotlight: Demonstrating Systematic Endobronchial Ultrasound to New Endoscopists
Published on: August 11, 2023
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021