Related Experiment Video
Updated: Jul 7, 2026

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Effect of low-complexity regions on protein structure determination.
Ryan M Bannen1, Craig A Bingman, George N Phillips
1Department of Biochemistry, University of Wisconsin-Madison, 433 Babcock Drive, Madison, WI 53711, USA.
Journal of Structural and Functional Genomics
|February 28, 2008
Summary
Low-complexity protein sequences are abundant but rarely found in the Protein Data Bank. This study analyzes trends in protein structure determination targets, offering insights for selecting future research candidates.
Area of Science:
- Biochemistry
- Structural Biology
- Bioinformatics
Background:
- Protein sequences with quasi-repetitive amino acid arrangements, termed low-complexity sequences, are prevalent in genomic and proteomic databases like Swiss-Prot.
- These low-complexity sequences are notably under-represented in the structure-based Protein Data Bank (PDB).
- Structural genomics initiatives have historically used the absence of low-complexity sequences as a criterion for selecting protein targets with a high likelihood of successful structure determination.
Purpose of the Study:
- To investigate trends related to low-complexity sequences in protein structure determination.
- To evaluate the utility of low-complexity sequences in the context of target selection for structural genomics.
- To provide data-driven insights for refining target selection strategies in structural biology.
Main Methods:
- Analysis of protein sequence data from the Protein Data Bank (PDB).
- Examination of data from structural genomics databases, including TargetDB and PepcDB.
- Comparative analysis of sequence characteristics and structure determination success rates.
Main Results:
- Identification of specific trends concerning the presence or absence of low-complexity sequences in proteins targeted for structure determination.
- Quantification of the under-representation of low-complexity sequences within the PDB.
- Observation of patterns in target selection by structural genomics groups concerning these sequences.
Conclusions:
- The findings offer valuable insights into the relationship between low-complexity sequences and protein structure determination.
- The study suggests that current target selection strategies may benefit from a nuanced consideration of low-complexity sequences.
- Recommendations are provided for optimizing the selection process for structural genomics projects based on sequence characteristics.
Related Concept Videos
Protein Folding
Overview
Protein Folding
Proteins are chains of amino acids linked together by peptide bonds. Upon synthesis, a protein folds into a three-dimensional conformation, critical to its biological function. Interactions between its constituent amino acids guide protein folding, and hence the protein structure is primarily dependent on its amino acid sequence.
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Protein Organization
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence.
The primary structure of a protein is its amino acid sequence.
Conservation of Protein Domains
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...

