Related Experiment Video
Updated: Apr 24, 2026

Annotation of Plant Gene Function via Combined Genomics, Metabolomics and Informatics
Published on: June 17, 2012
Integrating information retrieval with distant supervision for gene ontology annotation
Dongqing Zhu1, Dingcheng Li2, Ben Carterette2
1Department of Health Sciences Research, Mayo Clinic, 200 First St SW, Rochester, MN 55905 and Department of Computer & Information Sciences, University of Delaware, 101 SMITH HALL, Newark, DE 19716, USA Department of Health Sciences Research, Mayo Clinic, 200 First St SW, Rochester, MN 55905 and Department of Computer & Information Sciences, University of Delaware, 101 SMITH HALL, Newark, DE 19716, USA.
Researchers developed systems for the Gene Ontology Curation task, improving identification of gene evidence sentences and prediction of gene ontology terms from scientific articles.
Area of Science:
- Bioinformatics
- Computational Biology
- Biomedical Informatics
Background:
- The Gene Ontology (GO) provides a standardized vocabulary for gene function annotation.
- Accurate GO term assignment is crucial for understanding gene roles in biological processes.
- Automating GO curation from full-text articles presents significant challenges.
Purpose of the Study:
- To participate in the BioCreative IV Gene Ontology Curation task.
- To develop and evaluate systems for identifying GO evidence sentences (GOESs).
- To develop and evaluate systems for predicting GO terms for genes.
Main Methods:
- For GOES identification (Subtask A): Logistic regression model trained on existing annotations, supplemented with external negative data, followed by a greedy gene-sentence association approach.
- For GO term prediction (Subtask B): Two system types were developed: (i) search-based systems using information retrieval and distant supervision on various text granularities, and (ii) a similarity-based system measuring word distance to GO terms/synonyms.
Main Results:
- Subtask A best system achieved an F1 score of 0.27 (exact match) and 0.387 (relaxed overlap match).
- Subtask B best system (search-based) achieved an F1 score of 0.075 (exact match) and 0.301 (hierarchical match).
- Search-based systems for Subtask B significantly outperformed the similarity-based system.
Conclusions:
- The developed systems demonstrate progress in automated GO curation tasks.
- Search-based approaches show promise for predicting GO terms from literature.
- Further refinement is needed to improve performance, particularly for exact match scenarios.
Related Concept Videos
Genome Annotation and Assembly
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
Organization of Genes
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Gene Families

