Related Experiment Video
Updated: Jul 21, 2025

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
68.7K
Petascale Homology Search for Structure Prediction
Sewon Lee1, Gyuri Kim1, Eli Levy Karin2
1School of Biological Sciences, Seoul National University, Seoul 08826, South Korea.
Biorxiv : the Preprint Server for Biology
|July 28, 2023
Summary
Leveraging the Sequence Read Archive (SRA) for multiple sequence alignments (MSAs) significantly improved protein structure prediction accuracy. This enhanced MSA data boosted AlphaFold2-based predictions, advancing computational biology methods.
Area of Science:
- Computational Biology
- Structural Biology
- Bioinformatics
Background:
- Multiple sequence alignments (MSAs) are crucial for protein structure prediction, as evidenced by AlphaFold2's success in the CASP15 competition.
- The effective utilization of MSAs remains a key area for advancing prediction accuracy.
Conclusions:
- Large-scale SRA data significantly enhances MSA quality and protein structure prediction accuracy.
- Combining SRA-derived homologs with advanced computational strategies offers a powerful approach to improve protein structure modeling.
- These findings provide valuable insights for optimizing MSA generation in protein structure prediction pipelines.
Related Concept Videos
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K
Protein Organization
6.6K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.6K
Globular and Fibrous Proteins
43.8K
Many proteins can be classified into two distinct subtypes - globular or fibrous. These two types differ in their shapes and solubilities.
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
Globular proteins are also known as spheroproteins and typically are approximately round in shape. They contain a mix of amino acid types and contain differing sequences in their primary structures. Globular proteins have many different functions, such as enzymes, cellular messengers, and molecular transporters. These roles often require the proteins to be...
43.8K
Gene Families
8.9K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.9K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K

