Related Experiment Video
Updated: Jul 7, 2026

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
A template-finding algorithm and a comprehensive benchmark for homology modeling of proteins
Brinda Kizhakke Vallat1, Jaroslaw Pillardy, Ron Elber
1Department of Computer Science, Cornell University, Upson Hall 4130, Ithaca, New York 14853, USA.
Proteins
|February 27, 2008
Summary
This study introduces a novel decision tree method for identifying accurate protein templates in homology modeling. The approach significantly improves template detection accuracy compared to existing methods like PSI-BLAST.
Area of Science:
- Computational biology
- Structural bioinformatics
- Protein structure prediction
Background:
- Homology modeling is crucial for protein structure prediction, relying on identifying suitable template structures.
- Accurate template identification is a critical bottleneck in the homology modeling pipeline.
Purpose of the Study:
- To develop an improved method for identifying reliable protein templates for homology modeling.
- To enhance the accuracy and efficiency of template selection using a machine learning approach.
Main Methods:
- A large-scale dataset of protein sequence pairs was constructed from the Protein Data Bank (PDB).
- A decision tree classifier was trained using discriminatory learning on millions of true and false template matches.
- Each decision tree branch incorporated a mathematical programming model for refined discrimination.
Main Results:
- The developed decision tree model demonstrated significant enrichment of true templates (50-100%) over PSI-BLAST on independent PDB sets and CASP7 sequences.
- Verification using Modeller confirmed that high TM structural-alignment scores correlate with the generation of acceptable atomically detailed structural models.
Conclusions:
- The novel decision tree approach offers a substantial improvement in identifying correct protein templates for homology modeling.
- This method enhances the reliability of structural model generation by improving the quality of template selection.
Related Concept Videos
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Protein Organization
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence.
The primary structure of a protein is its amino acid sequence.
Protein Organization
Overview
Protein Organization
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence.
The primary structure of a protein is its amino acid sequence.
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...

