Related Experiment Video
Updated: Apr 25, 2026

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
The impact of incomplete knowledge on the evaluation of protein function prediction: a structured-output learning
Yuxiang Jiang1, Wyatt T Clark1, Iddo Friedberg2
1Department of Computer Science and Informatics, Indiana University, Bloomington, IN, USA, Department of Microbiology and Department of Computer Science and Software Engineering, Miami University, Oxford, OH, USA.
Motivation:
The automated functional annotation of biological macromolecules is a problem of computational assignment of biological concepts or ontological terms to genes and gene products. A number of methods have been developed to computationally annotate genes using standardized nomenclature such as Gene Ontology (GO). However, questions remain about the possibility for development of accurate methods that can integrate disparate molecular data as well as about an unbiased evaluation of these methods. One important concern is that experimental annotations of proteins are incomplete. This raises questions as to whether and to what degree currently available data can be reliably used to train computational models and estimate their performance accuracy.
Results:
We study the effect of incomplete experimental annotations on the reliability of performance evaluation in protein function prediction. Using the structured-output learning framework, we provide theoretical analyses and carry out simulations to characterize the effect of growing experimental annotations on the correctness and stability of performance estimates corresponding to different types of methods. We then analyze real biological data by simulating the prediction, evaluation and subsequent re-evaluation (after additional experimental annotations become available) of GO term predictions. Our results agree with previous observations that incomplete and accumulating experimental annotations have the potential to significantly impact accuracy assessments. We find that their influence reflects a complex interplay between the prediction algorithm, performance metric and underlying ontology. However, using the available experimental data and under realistic assumptions, our results also suggest that current large-scale evaluations are meaningful and almost surprisingly reliable.
Supplementary Information:
Supplementary data are available at Bioinformatics online.
More Related Videos
07:35A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
Related Concept Videos
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein Organization
The primary structure of a protein is its amino acid sequence....
Protein Organization
Protein Organization
Protein Organization