Related Experiment Video
Updated: Jun 30, 2026

Assessment of Immunologically Relevant Dynamic Tertiary Structural Features of the HIV-1 V3 Loop Crown R2 Sequence by ab initio Folding
Published on: September 15, 2010
Folding the unfoldable 2: using AlphaFold and ESMFold to explore spurious proteins
1European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Hinxton, CB10 1SD, UK.
Newer protein structure prediction tools like AlphaFold3 and ESMFold still predict short spurious sequences with high confidence. Discrimination improves for longer sequences, aiding in identifying genuine proteins and potential errors.
Area of Science:
- Computational biology
- Structural bioinformatics
- Protein structure prediction
Background:
- Spurious protein sequences from gene prediction errors are theoretically unlikely to form stable structures.
- Previous studies indicated AlphaFold2 inaccurately assigned high confidence scores (pLDDT) to short spurious sequences, hindering distinction from real proteins.
- The ability of advanced structure prediction tools to differentiate between genuine and artifactual protein sequences remains a critical question.
Purpose of the Study:
- To evaluate if newer protein structure prediction methods, ESMFold and AlphaFold3, can better distinguish between spurious and genuine protein sequences compared to AlphaFold2.
- To assess the performance of these methods in assigning confidence scores (pLDDT) to short spurious sequences.
- To explore the potential of structure prediction scores for identifying spurious proteins at scale.
Main Methods:
- Comparative analysis of protein structure prediction by ESMFold, AlphaFold2, and AlphaFold3 on known spurious sequences (AntiFam).
- Evaluation of prediction confidence scores (pLDDT and pTM) for sequences of varying lengths.
- Development and application of a Gaussian Process Model using structure prediction scores to identify potential spurious proteins in large datasets (AlphaFold DB).
Main Results:
- ESMFold, AlphaFold2, and AlphaFold3 all assigned high pLDDT scores to short spurious sequences.
- Improved discrimination between spurious and real proteins was observed for sequences exceeding 100 amino acids.
- Analysis of disparate pLDDT and pTM scores led to the identification of potentially novel spurious ORFs and a possibly non-spurious AntiFam entry.
- A Gaussian Process Model utilizing structure prediction scores showed promise in identifying potential spurious proteins at scale when combined with other methods.
Conclusions:
- Current advanced protein structure prediction tools, including AlphaFold3 and ESMFold, continue to assign high confidence to short spurious sequences.
- Sequence length is a critical factor, with better discrimination achieved for proteins over 100 amino acids.
- Structure prediction scores offer a valuable, albeit complementary, approach for identifying spurious proteins, especially when integrated into larger identification pipelines.
Related Concept Videos
Protein Folding
Protein Folding
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
Protein Folding
Molecular Chaperones and Protein Folding
The...
Molecular Chaperones and Protein Folding
The...
Amyloid Fibrils
Amyloid deposits were observed as early as 1639 in the liver and the spleen. In 1854, Rudolph Virchow performed iodine staining, normally used to...

