Related Experiment Video
Updated: Jun 25, 2025

A Data Integration Workflow to Identify Drug Combinations Targeting Synthetic Lethal Interactions
Published on: May 27, 2021
DUVEL: an active-learning annotated biomedical corpus for the recognition of oligogenic combinations
Charlotte Nachtegael1,2, Jacopo De Stefani2, Anthony Cnudde2,3
1Interuniversity Institute of Bioinformatics in Brussels, Université Libre de Bruxelles-Vrije Universiteit Brussel, Boulevard du Triomphe, CP 263, Brussels 1050, Belgium.
This study introduces DUVEL, a novel dataset for extracting oligogenic variant combinations crucial for understanding complex diseases. Active learning efficiently created this dataset, improving biomedical relation extraction model performance.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Biomedical relation extraction (bioRE) datasets currently focus on single variants, limiting research into complex genetic interactions.
- Understanding digenic and oligogenic variant combinations is vital for elucidating disease etiologies and epistatic effects.
- Existing literature highlights the importance of multi-variant interactions, yet lacks dedicated datasets for computational analysis.
Purpose of the Study:
- To create a unique dataset (DUVEL) for training computational tools to extract oligogenic variant combinations from scientific literature.
- To address the scarcity of datasets for multi-gene variant interactions in biomedical relation extraction.
- To leverage active learning to optimize the annotation process for efficiency and cost-effectiveness.
Main Methods:
- Utilized active learning (AL) to guide the annotation of text fragments containing potential digenic variant combinations (gene-variant-gene-variant).
- Pre-annotated 85 full-text articles from the Oligogenic Diseases Database (OLIDA) using PubTator.
- Annotated extracted text fragments using ALAMBIC, an AL-based annotation platform, to create the DUVEL dataset.
Main Results:
- The DUVEL dataset comprises 8442 text fragments, with 794 positive instances of oligogenic variant relations, covering 95% of annotated articles.
- Fine-tuning state-of-the-art biomedical language models (BiomedBERT, BiomedBERT-large, BioLinkBERT, BioM-BERT) on DUVEL significantly improved performance.
- BiomedBERT-large achieved the highest F1 score of 0.84 for gene-variant pair detection after fine-tuning, demonstrating the dataset's utility.
Conclusions:
- The DUVEL dataset provides a valuable resource for advancing biomedical relation extraction, specifically for 4-ary relations between two genes and two variants.
- Active learning is an effective strategy for creating specialized bioRE datasets, reducing annotation costs and improving efficiency.
- The DUVEL dataset is freely available, facilitating research in computational biology and the curation of scientific literature on complex genetic interactions.
Related Concept Videos
Multiple Allele Traits
Complementary DNA
Dihybrid Crosses
Genetic Lingo
Pedigree Analysis
Pleiotropy

