Related Experiment Video
Updated: Feb 1, 2026

08:09
Annotation of Plant Gene Function via Combined Genomics, Metabolomics and Informatics
Published on: June 17, 2012
20.6K
Automatic gene annotation using GO terms from cellular component domain.
Ruoyao Ding1, Yingying Qu2, Cathy H Wu3
1School of Information Science and Technology, Guangdong University of Foreign Studies, Guangzhou, China.
BMC Medical Informatics and Decision Making
|December 12, 2018
Summary
This study introduces a novel relation extraction approach to automate Gene Ontology (GO) annotation for cellular components, significantly speeding up the process for bio-annotators.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Gene Ontology (GO) provides structured information on gene product functions across Cellular Component (CC), Molecular Function (MF), and Biological Process (BP) domains.
- Manual GO annotation, assigning GO terms to genes from literature, is time-consuming and labor-intensive for Model Organism Databases (MODs).
- Current manual annotation efforts can only cover a fraction of relevant scientific articles due to resource limitations.
Purpose of the Study:
- To develop an automated approach for gene annotation using GO terms, specifically within the Cellular Component (CC) domain.
- To address the challenge of efficiently assigning subcellular location and protein complex information to genes.
Main Methods:
- The study frames gene annotation with CC domain GO terms as two relation extraction subtasks: identifying protein subcellular locations and protein complex subunits.
- A relation extraction approach utilizing triggers and syntactic dependencies was employed for both subtasks.
- The method was evaluated on the publicly available BC4GO test corpus.
Main Results:
- The proposed approach achieved a F1-score of 71% for predicting GO terms within the CC domain for given genes.
- The system demonstrated a precision of 91% and a recall of 58% in these predictions.
- These results indicate a significant improvement in the efficiency and accuracy of automated GO annotation.
Conclusions:
- A novel method was developed to automate gene annotation with CC domain GO terms by treating it as two relation extraction tasks.
- The approach successfully accelerates the gene annotation process for bio-annotators.
- The evaluation results confirm the effectiveness of this automated strategy for GO annotation.
Related Concept Videos
Automatic Processing and Automatic Social Behavior
252
Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
252
Genome Annotation and Assembly
21.0K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
21.0K
Conservation of Protein Domains Over Different Proteins
14.5K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.5K
Long-term Depression
33.3K
Long-term depression, or LTD, is one of the ways by which synaptic plasticity—changes in the strength of chemical synapses—can occur in the brain. LTD is the process of synaptic weakening that occurs over time between pre and postsynaptic neuronal connections. The synaptic weakening of LTD works in opposition to synaptic strengthening by long-term potentiation (LTP) and together are the main mechanisms that underlie learning and memory.
33.3K
Components of Stress
536
Stress analysis under multiple loading conditions is intricate, necessitating a comprehensive grasp of normal and shearing stresses. Consider a small cube at point O, subjected to stress on all six faces, visible or not. Normal stress components σx, σy, σz act perpendicularly to the x, y, and z axes. Shearing stress components τxy and τxz are exerted on faces perpendicular to these axes.
Interestingly, the hidden cube faces also experience these stresses, equal and...
Interestingly, the hidden cube faces also experience these stresses, equal and...
536
Components of Language
823
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
823

