Related Experiment Video
Updated: Apr 30, 2026

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
Redundancy-weighting for better inference of protein structural features
Chen Yanover1, Natalia Vanetik1, Michael Levitt1
1Machine Learning for Healthcare and Life-Sciences, Analytics Department, IBM Research Laboratory, Haifa, 3490002, Department of Software Engineering, Shamoon College of Engineering, Beer-Sheva 84100, Israel, Department of Structural Biology, Stanford University School of Medicine, Stanford, CA 94305, USA, Department of Computer Science, University of Haifa, Mount Carmel, Haifa, 3498838 and Departments of Life Sciences and Computer Science, Ben-Gurion University of the Negev, Beer-Sheva, 84105, Israel.
Redundancy-weighted datasets, which account for protein structure diversity, offer smoother, more accurate distributions than non-redundant sets. This approach enhances protein structure prediction and knowledge-based potentials.
Area of Science:
- Structural bioinformatics
- Computational biology
- Protein structure analysis
Background:
- Protein Data Bank (PDB) data is biased, with redundant entries and underrepresented protein families.
- Non-redundant PDB subsets mask sequence-structure relationships and limit the use of new structural data.
- Existing methods struggle to capture the full diversity of protein sequences and conformations.
Purpose of the Study:
- To systematically compare redundancy-weighted datasets with traditional non-redundant datasets.
- To evaluate the impact of different weighting schemes on structural feature distributions.
- To demonstrate the advantages of redundancy-weighting for protein structure analysis.
Main Methods:
- Exploration of redundancy-weighted datasets, assigning weights inversely proportional to homolog counts.
- Systematic comparison of three weighting schemes against non-redundant datasets.
- Analysis of structural feature distributions and their properties (smoothness, robustness, correctness).
Main Results:
- Redundancy-weighted datasets produce smoother (higher entropy) distributions of structural features compared to non-redundant datasets.
- The distributions derived from redundancy-weighted sets are shown to be more robust and accurate.
- Smoothed distributions offer a more comprehensive view of protein structural diversity.
Conclusions:
- Redundancy-weighting provides a more accurate representation of protein structural space than non-redundant datasets.
- Improved distributions can enhance the accuracy of knowledge-based potentials.
- This approach has the potential to significantly improve protein structure prediction methods and model-driven molecular biology.
Related Concept Videos
Protein Folding Quality Check in the RER
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Protein Folding
Protein Folding
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
Protein Folding

