Related Experiment Video
Updated: Sep 4, 2025

Biosensor-based High Throughput Biopanning and Bioinformatics Analysis Strategy for the Global Validation of Drug-protein Interactions
Published on: December 1, 2020
A sequence labeling framework for extracting drug-protein relations from biomedical literature
Ling Luo1, Po-Ting Lai1, Chih-Hsuan Wei1
1National Center for Biotechnology Information (NCBI), National Library of Medicine (NLM), National Institutes of Health (NIH), 8600 Rockville Pike, Bethesda, MD 20894, USA.
This study introduces a sequence labeling framework for extracting drug-protein interactions, outperforming traditional text classification. Ensemble methods further boosted performance, achieving a 0.800 F1-score in the BioCreative VII challenge.
Area of Science:
- Biomedical Natural Language Processing
- Computational Biology
- Drug Discovery Informatics
Background:
- Automatic extraction of drug-protein interactions is crucial for drug discovery, repurposing, and biomedical knowledge graph construction.
- The BioCreative VII challenge's DrugProt track aimed to advance drug-protein relation extraction.
Purpose of the Study:
- To develop and evaluate novel frameworks for drug-protein relation extraction.
- To compare the efficacy of sequence labeling versus text classification approaches.
- To improve performance through ensemble methods and advanced language models.
Main Methods:
- Comparison of cutting-edge biomedical pre-trained language models within text classification and sequence labeling frameworks.
- Exploration of ensemble methods, including majority voting, to combine model predictions.
- Implementation of a sequence labeling framework for drug-protein relation extraction.
Main Results:
- The sequence labeling framework demonstrated superior efficiency and performance compared to the text classification framework.
- An ensemble of models using majority voting achieved an F1-score of 0.795 on the official test set.
- The best performing approach, an ensemble of sequence labeling models, reached an F1-score of 0.800.
Conclusions:
- Sequence labeling is a more effective framework for drug-protein relation extraction.
- Ensemble methods significantly enhance the performance of relation extraction models.
- The developed approach achieved state-of-the-art results in the BioCreative VII DrugProt challenge.
More Related Videos
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-protein Interfaces
Protein-Drug Binding: Determination Methods
Indirect methods involve isolating the bound drug from its free form in biological samples such as blood, serum, or plasma. These techniques aim to measure the percentage of drugs bound to proteins. Equilibrium dialysis is a commonly used method where the free drug concentration at equilibrium is measured by separating the bound...
Ligand Binding and Linkage
Protein-Drug Binding: Mechanism and Kinetics
Various forces drive these interactions, including hydrogen bonds, hydrophobic interactions, ionic bonds, electrostatic interactions, and van der Waals forces. These bonds enable drugs to bind to specific sites on proteins,...

