Related Experiment Video
Updated: Jul 7, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Discrimination between distant homologs and structural analogs: lessons from manually constructed, reliable data
Hua Cheng1, Bong-Hyun Kim, Nick V Grishin
1Howard Hughes Medical Institute, University of Texas Southwestern Medical Center, 5323 Harry Hines Boulevard, Dallas, TX 75390-9050, USA. hua.cheng@utsouthwestern.edu
This study develops a machine learning classifier to distinguish protein homologs from analogs using evolutionary and structural data. Profile sequence scores from structural alignments best differentiate these protein relationships, aiding evolutionary studies.
Area of Science:
- Computational Biology
- Structural Bioinformatics
- Evolutionary Biology
Background:
- Understanding protein sequence, structure, and function is aided by evolutionary context.
- Homologs share similarities due to common ancestry, while analogs converge on similar structures.
- Previous work established reliable databases for homologs and analogs.
Purpose of the Study:
- To compare existing homolog and analog datasets.
- To develop a machine learning classifier, specifically a support vector machine (SVM), to discriminate between homologs and analogs.
- To identify the most effective features for distinguishing remote homologs from structural analogs.
Main Methods:
- Comparison of two previously assembled databases of protein homologs and analogs.
- Development of a support vector machine (SVM) classifier utilizing various similarity scores.
- Evaluation of sequence and structure-based similarity scores, with a focus on profile sequence scores derived from structural alignments.
Main Results:
- Profile sequence scores, computed from structural alignments, demonstrated superior performance in discriminating between remote homologs and structural analogs compared to general sequence or structure scores.
- The SVM classifier successfully identified 76% of remote homologs within the Structural Classification of Proteins (SCOP) database (domains in the same superfamily but different families).
- Novel homologous relationships were identified between SCOP domains across different superfamilies, folds, and even classes.
Conclusions:
- Profile sequence scores derived from structural alignments are highly effective discriminators for remote homologs and structural analogs.
- The developed SVM classifier accurately identifies remote homologous relationships within curated protein databases like SCOP.
- This approach reveals deeper evolutionary connections between proteins, extending beyond traditional classifications.
More Related Videos
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025
08:04Identification and Classification of Position-specific GABAA Receptor Subunit Missense Variants for Their Role In Hippocampal Pyramidal Neurons
Published on: June 6, 2025
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Protein Families
Gene Families
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Convergent Evolution
¹H NMR Chemical Shift Equivalence: Homotopic and Heterotopic Protons