Related Experiment Video
Updated: Aug 14, 2026

12:00
A Practical Guide to Phylogenetics for Nonexperts
Published on: February 5, 2014
Automated methods of predicting the function of biological sequences using GO and BLAST
Craig E Jones1, Ute Baumann, Alfred L Brown
1Australian Centre for Plant Functional Genomics, University of Adelaide, South Australia, 5064, Australia. craig@cs.adelaide.edu.au
BMC Bioinformatics
|November 18, 2005
Summary
Automated biological function prediction for genomic sequences is crucial. Accuracy benchmarking effectively compares prediction methods, with data mining yielding the most accurate results for Gene Ontology (GO) annotation.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Increasing genomic data necessitates automated, accurate methods for predicting biological functions of novel sequences.
- Gene Ontology (GO) provides a structured vocabulary for annotating sequences with biological context.
- Evaluating competing function prediction designs requires robust accuracy benchmarking.
Purpose of the Study:
- To demonstrate accuracy benchmarking for evaluating biological function predictors.
- To assess the impact of term distance on accuracy scores.
- To compare various sequence annotation methods using accuracy benchmarking.
Main Methods:
- Utilized Gene Ontology (GO) for functional annotation of sequences.
- Employed Basic Local Alignment Search Tool (BLAST) to find similar annotated sequences.
- Developed and compared several annotation methods based on terms from top BLAST matches against a 'best BLAST' benchmark.
Main Results:
- Precision and recall increase with allowed distance between predicted and true GO terms.
- Accuracy benchmarking effectively compares annotation methods.
- A discriminant function method showed superior precision and recall compared to other tested approaches, including the 'best BLAST' method.
Conclusions:
- Requiring exact term matches for accuracy improves reliability over methods allowing related terms.
- Accuracy benchmarking is effective for comparing BLAST-based GO annotator designs.
- Data mining techniques produced the most accurate annotation method, highlighting its utility for developing high-quality annotators.
Related Concept Videos
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...

