Related Experiment Video
Updated: Jun 13, 2025

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
68.6K
The simplicity of protein sequence-function relationships
Yeonwoo Park1,2, Brian P H Metzger3,4, Joseph W Thornton5,6
1Committee on Genetics, Genomics, and Systems Biology, University of Chicago, Chicago, IL, USA.
Nature Communications
|September 11, 2024
Summary
Protein sequence-function relationships are simpler than previously thought. A new reference-free method reveals that basic amino acid effects and pairwise interactions explain most protein function, challenging complex epistasis models.
Area of Science:
- Molecular Biology
- Genetics
- Biophysics
Background:
- High-order epistatic interactions are believed to govern protein sequence-function relationships, implying complexity and unpredictability.
- Previous studies may overestimate epistasis due to reference-dependent analyses and failure to account for global nonlinearities.
Purpose of the Study:
- To develop and validate a reference-free method for inferring protein sequence-function relationships.
- To determine the true extent of epistasis and identify key determinants of protein function.
Main Methods:
- A novel reference-free computational method was developed to jointly infer epistatic interactions and global nonlinearity.
- The method was applied to 20 diverse experimental datasets covering protein sequence-function relationships.
Main Results:
- Amino acid effects and pairwise interactions, with a simple nonlinearity, explain a median of 96% of phenotypic variance across datasets.
- Higher-order epistasis affects only a small fraction of genotypes, and sequence-function relationships are sparse.
- The new method is robust to noise, missing data, and model misspecification.
Conclusions:
- Protein sequence-function causality is largely simple and predictable, dominated by context-independent effects and pairwise interactions.
- This simplification opens avenues for tractable methods to characterize protein genetic architecture.
- The findings challenge the pervasive view of complex, high-order epistasis in protein evolution.
More Related Videos
Related Concept Videos
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Protein Folding
117.7K
Overview
117.7K
Protein and Protein Structure
79.1K
Proteins are one of the most abundant organic molecules in living systems and have the most diverse range of functions of all macromolecules. Proteins may be structural, regulatory, contractile, or protective. They may serve in transport, storage, or membranes; or they may be toxins or enzymes. Their structures, like their functions, vary greatly. They are all, however, amino acid polymers arranged in a linear sequence.
A protein's shape is critical to its function. For example, an enzyme...
A protein's shape is critical to its function. For example, an enzyme...
79.1K
Protein Organization
6.3K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.3K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K

