Related Experiment Video
Updated: Sep 4, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A BERT-based ensemble learning approach for the BioCreative VII challenges: full-text chemical identification and
Sheng-Jie Lin1, Wen-Chao Yeh2, Yu-Wen Chiu1
1Graduate Institute of Data Science, Taipei Medical University, No. 172-1, Section 2, Keelung Rd, Dáan District, Taipei City 106, Taiwan.
This study enhances biomedical text analysis using Bidirectional Encoder Representations from Transformers (BERT) models. The research achieves state-of-the-art results in identifying chemical entities and COVID-19 literature topics.
Area of Science:
- Biomedical Natural Language Processing
- Machine Learning in Bioinformatics
- Computational Chemistry
Background:
- Biomedical text mining requires advanced models for accurate information extraction.
- Existing Bidirectional Encoder Representations from Transformers (BERT) models offer potential but require domain-specific optimization.
- The BioCreative VII Challenge presents specific tracks (NLM CHEM, LitCovid) for evaluating biomedical NLP systems.
Purpose of the Study:
- To evaluate state-of-the-art biomedical-specific BERT models for the NLM CHEM and LitCovid tracks.
- To propose and validate a BERT-based ensemble learning approach for improved performance.
- To assess the effectiveness of a novel Medical Subject Headings identifier (MeSH ID) normalization algorithm.
Main Methods:
- Exploration and ensemble of various biomedical-specific pre-trained BERT models.
- Application of the ensemble approach to the NLM CHEM and LitCovid tracks of the BioCreative VII Challenge.
- Development and evaluation of a Medical Subject Headings identifier (MeSH ID) normalization algorithm for entity normalization.
Main Results:
- Achieved F1-scores of 85% (strict) and 91.8% (approximate) on the NLM-CHEM track.
- The MeSH ID normalization algorithm yielded F1-scores of approximately 80% in both strict and approximate evaluations.
- Demonstrated state-of-the-art performance on the LitCovid track for COVID-19 literature topic detection, outperforming existing methods.
Conclusions:
- The proposed BERT-based ensemble learning approach significantly improves performance in biomedical text analysis tasks.
- The MeSH ID normalization algorithm is effective for accurate entity normalization in biomedical contexts.
- The methodology establishes a new benchmark for topic detection in COVID-19 literature.
More Related Videos
07:11Author Spotlight: Emerging Technologies and Advanced Tools for Decoding Metabolomics Data Analysis
Published on: November 10, 2023
13:49Semi-automated Biopanning of Bacterial Display Libraries for Peptide Affinity Reagent Discovery and Analysis of Resulting Isolates
Published on: December 6, 2017
Related Concept Videos
Classification of Titrimetric Analysis Based on Reaction Types
Titrations between an acid and a base lead to neutralization reactions that form...
Methods of Classification and Identification
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
MALDI-TOF Mass Spectrometry
Matrix-assisted laser desorption ionization (MALDI) is a commonly...
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...