Related Experiment Video
Updated: Sep 28, 2025

07:35
A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
1.8K
TBGA: a large-scale Gene-Disease Association dataset for Biomedical Relation Extraction
Stefano Marchesin1, Gianmaria Silvello2
1Department of Information Engineering, University of Padova, Padova, Italy. stefano.marchesin@unipd.it.
BMC Bioinformatics
|April 1, 2022
Summary
We created TBGA, a large dataset for extracting gene-disease associations (GDAs) from scientific literature. This resource aids in training models for biomedical relation extraction, advancing GDA discovery.
Area of Science:
- Biomedical informatics
- Computational biology
- Bioinformatics
Background:
- Biomedical databases require extensive manual updates.
- Biomedical Relation Extraction (BioRE) automates data population.
- Gene-Disease Association (GDA) extraction is a key BioRE task with limited training data.
Purpose of the Study:
- To develop a large-scale dataset for training GDA extraction models.
- To address the scarcity of resources for GDA extraction model development.
Main Methods:
- Leveraged the DisGeNET database for data curation.
- Semi-automatically annotated a dataset from over 700K publications.
- Constructed the TBGA dataset with over 200K instances and 100K gene-disease pairs.
Main Results:
- Developed TBGA, a comprehensive dataset for GDA extraction.
- TBGA comprises sentences, extracted GDAs, and gene-disease pair information.
- The dataset is one of the largest available for GDA extraction.
Conclusions:
- TBGA is a large and challenging dataset suitable for GDA extraction.
- Evaluated state-of-the-art models on TBGA, demonstrating its utility.
- The publicly available TBGA dataset aims to advance BioRE model development for GDA extraction.
Related Concept Videos
Genome-wide Association Studies-GWAS
14.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.5K
Genomics
37.8K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
37.8K
Genome Annotation and Assembly
19.4K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.4K

