Related Experiment Video
Updated: Oct 12, 2025

08:01
A Web Tool for Generating High Quality Machine-readable Biological Pathways
Published on: February 8, 2017
17.9K
Automatic consistency assurance for literature-based gene ontology annotation
Jiyu Chen1, Nicholas Geard1, Justin Zobel1
1School of Computing and Information Systems, University of Melbourne, Melbourne, 3010, Australia.
BMC Bioinformatics
|November 26, 2021
Summary
We developed a text mining method to automatically detect inconsistencies in gene ontology (GO) annotations, improving the accuracy of biological database curation and ensuring data reliability.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Literature-based gene ontology (GO) annotation is crucial for understanding gene functions but manual consistency checks are time-consuming.
- Maintaining GO annotation consistency with evolving literature and vocabulary is a significant challenge in bioinformatics.
- Automated methods are needed to ensure the reliability of GO annotations.
Purpose of the Study:
- To formalize biological database annotation inconsistencies and identify distinct types.
- To develop an efficient text mining method for automatically distinguishing consistent from inconsistent GO annotations.
- To evaluate the performance and robustness of the proposed method.
Main Methods:
- Formalization of four distinct types of biological database annotation inconsistencies.
- Application of state-of-the-art text mining models for automated inconsistency detection.
- Evaluation using a synthetic dataset (BC4GO) with manipulated annotations and detailed error analysis.
Main Results:
- The proposed method accurately distinguishes between consistent and inconsistent GO annotations.
- High precision was achieved on confident predictions, validated through error analysis.
- Two models demonstrated robustness against updates in the GO vocabulary.
Conclusions:
- The developed text mining approach effectively identifies GO annotation inconsistencies.
- The method shows significant value in supporting human-in-the-loop curation processes.
- This approach enhances the reliability and maintainability of biological databases.
Related Concept Videos
Genome Annotation and Assembly
19.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.5K
Gene Families
9.3K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
9.3K
Genome-wide Association Studies-GWAS
14.7K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.7K
Organization of Genes
71.0K
Overview
71.0K
Ligand Binding and Linkage
5.1K
Allosteric proteins have more than one ligand binding site; the binding of a ligand to any of these sites influences the binding of ligands to the other sites. When a protein is allosteric, its binding sites are called coupled or linked. In the case of enzymes, the site that binds to the substrate is known as the active site and the other site is known as the regulatory site. When a ligand binds to the regulatory site, this leads to conformational changes in the protein that can influence...
5.1K
Genetic Lingo
106.7K
Overview
106.7K

