Related Experiment Video
Updated: Jul 4, 2025

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
68.7K
Petabase-Scale Homology Search for Structure Prediction
Sewon Lee1, Gyuri Kim1, Eli Levy Karin2
1School of Biological Sciences, Seoul National University, Gwanak-gu, Seoul 08826, South Korea.
Cold Spring Harbor Perspectives in Biology
|February 5, 2024
Summary
Leveraging the Sequence Read Archive (SRA) for protein structure prediction significantly boosts accuracy. Incorporating SRA data improved predictions for 66% of targets, outperforming standard methods.
Area of Science:
- Computational Biology
- Structural Biology
- Bioinformatics
Background:
- Multiple sequence alignments (MSAs) are crucial for protein structure prediction, as evidenced by AlphaFold2's success in the CASP15 competition.
- The effective utilization of MSAs remains a key area for advancing prediction accuracy.
Purpose of the Study:
- To investigate the impact of large-scale Sequence Read Archive (SRA) data on protein structure prediction accuracy.
- To evaluate the contribution of SRA-derived homologs and advanced ColabFold features to prediction performance.
Main Methods:
- A petabase-scale search of the SRA was performed to generate extensive aligned homologs for CASP15 targets.
- SRA-derived MSAs were merged with default ColabFold-search MSAs and used as input for ColabFold-predict.
- The effects of deep homology search and ColabFold's advanced features (e.g., increased recycles) were systematically tested.
Main Results:
- Using SRA data improved highly accurate predictions (GDT_TS > 70) to 66% for non-easy targets, compared to 52% with default ColabFold-search MSAs.
- The inclusion of SRA homologs was the primary factor in improving ColabFold's CASP15 ranking from 11th to 3rd place.
- Advanced ColabFold features also contributed to the overall improvement in prediction accuracy.
Conclusions:
- Expanding the scope of MSAs through large-scale data sources like the SRA is a powerful strategy for enhancing protein structure prediction.
- Combining SRA data with optimized computational approaches like ColabFold offers significant improvements in predicting protein structures.
- Further analysis of deep homology and advanced prediction features is warranted to fully leverage MSA data.
Related Concept Videos
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Organization
6.5K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.5K
Globular and Fibrous Proteins
43.7K
Many proteins can be classified into two distinct subtypes - globular or fibrous. These two types differ in their shapes and solubilities.
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
43.7K
Gene Families
8.8K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.8K

