Related Experiment Video
Updated: Apr 20, 2026

Annotation of Plant Gene Function via Combined Genomics, Metabolomics and Informatics
Published on: June 17, 2012
A robust data-driven approach for gene ontology annotation
1Department of Quantitative Health Sciences, University of Massachusetts Medical School, Worcester, MA, USA, Department of Computer Science, University of Massachusetts, Amherst, MA, USA and VA Central Western Massachusetts, Worcester, MA, USA liyanpeng.lyp@gmail.com.
Abstract:
Gene ontology (GO) and GO annotation are important resources for biological information management and knowledge discovery, but the speed of manual annotation became a major bottleneck of database curation. BioCreative IV GO annotation task aims to evaluate the performance of system that automatically assigns GO terms to genes based on the narrative sentences in biomedical literature. This article presents our work in this task as well as the experimental results after the competition. For the evidence sentence extraction subtask, we built a binary classifier to identify evidence sentences using reference distance estimator (RDE), a recently proposed semi-supervised learning method that learns new features from around 10 million unlabeled sentences, achieving an F1 of 19.3% in exact match and 32.5% in relaxed match. In the post-submission experiment, we obtained 22.1% and 35.7% F1 performance by incorporating bigram features in RDE learning. In both development and test sets, RDE-based method achieved over 20% relative improvement on F1 and AUC performance against classical supervised learning methods, e.g. support vector machine and logistic regression. For the GO term prediction subtask, we developed an information retrieval-based method to retrieve the GO term most relevant to each evidence sentence using a ranking function that combined cosine similarity and the frequency of GO terms in documents, and a filtering method based on high-level GO classes. The best performance of our submitted runs was 7.8% F1 and 22.2% hierarchy F1. We found that the incorporation of frequency information and hierarchy filtering substantially improved the performance. In the post-submission evaluation, we obtained a 10.6% F1 using a simpler setting. Overall, the experimental analysis showed our approaches were robust in both the two tasks.
Related Concept Videos
Genome Annotation and Assembly
Genomics
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Synthetic Biology
Golden rice
Golden rice is a genetically modified...
Gene Families
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...

