Related Experiment Video
Updated: Aug 8, 2026

12:00
A Practical Guide to Phylogenetics for Nonexperts
Published on: February 5, 2014
Evolutionary and structural feedback on selection of sequences for comparative analysis of proteins
I Mihalek1, I Res, O Lichtarge
1Department of Molecular and Human Genetics, Baylor College of Medicine, Houston, Texas, USA. imihalek@bcm.tmc.edu
Proteins
|January 7, 2006
Summary
Slowly evolving protein residues cluster in native folds and indicate functional surfaces. This study shows these two properties are linked, meaning one implies the other for protein analysis.
Area of Science:
- Biochemistry
- Structural Biology
- Bioinformatics
Background:
- Slowly evolving protein residues exhibit two key traits: spatial clustering in the native protein fold and delineation of functional surfaces involved in molecular interactions.
- These functional surfaces mediate protein interactions with other proteins or small molecules.
Purpose of the Study:
- To demonstrate the strong coupling between residue clustering and functional surface delineation in proteins.
- To establish a method for selecting homologous sequences to optimize the detection of functional surfaces.
Main Methods:
- Utilized multiple sequence alignment (MSA) related methods with careful selection of relevant sequences.
- Applied the methods to two distinct protein datasets: a small set of diverse proteins and a large set of homodimerizing enzymes.
Main Results:
- Confirmed that the clustering of slowly evolving residues in the native fold statistically implies their role in delineating functional surfaces.
- Showcased that the coupling between these two properties is strong enough for one to predict the other.
Conclusions:
- A simple rule for selecting homologous sequences for comparative protein analysis is proposed: select sequences such that observed residues at any evolutionary divergence level cluster on the folded protein.
- This approach optimizes the detection of potentially unknown functional surfaces in proteins.
Related Concept Videos
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Gene Evolution - Fast or Slow?
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...

