Generalizable and scalable protein stability prediction with rewired protein generative models
1School of Computational Science and Engineering, Georgia Institute of Technology, Atlanta, GA, USA.
Nature Communications
|December 20, 2025
Summary
We developed SPURS, a deep learning framework that enhances protein stability prediction by integrating generative models. This tool offers accurate, scalable predictions and broad applications in protein engineering and disease research.
Area of Science:
- Biochemistry
- Computational Biology
- Machine Learning
Background:
- Predicting protein thermostability changes from amino acid substitutions is crucial for disease research and protein engineering.
- Existing protein generative models show promise but their potential for stability prediction is underexploited.
Purpose of the Study:
- To present SPURS, a novel deep learning framework for accurate and efficient protein stability prediction.
- To leverage and integrate complementary protein generative models for enhanced performance.
- To enable broad applications in protein informatics and disease research.
Main Methods:
- SPURS integrates a protein language model and an inverse folding model.
- The unified framework is fine-tuned using large-scale thermostability data.
- Deep learning techniques are employed for stability prediction and analysis.
Main Results:
- SPURS achieves accurate, efficient, and scalable predictions of protein stability.
- The framework generalizes well to novel proteins and mutations.
- SPURS demonstrates utility in zero-shot functional residue identification and protein fitness prediction.
Conclusions:
- SPURS establishes a versatile tool for advancing protein stability prediction and engineering.
- The framework facilitates systematic analysis of stability-pathogenicity links in human diseases.
- SPURS enhances capabilities in protein informatics and practical protein design.
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
13.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
13.9K
Protein Folding Quality Check in the RER
5.0K
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...
5.0K
RNA Stability
35.6K
Intact DNA strands can be found in fossils, while scientists sometimes struggle to keep RNA intact under laboratory conditions. The structural variations between RNA and DNA underlie the differences in their stability and longevity. Because DNA is double-stranded, it is inherently more stable. The single-stranded structure of RNA is less stable but also more flexible and can form weak internal bonds. Additionally, most RNAs in the cell are relatively short, while DNA can be up to 250 million...
35.6K
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K
Conservation of Protein Domains
3.9K
3.9K
Protein Modifications in the RER
6.8K
Modification of secretory and transmembrane proteins entering the rough ER begins in the ER lumen. These modifications aid in protein folding and stabilize the acquired tertiary structure. Protein modifications in the rough ER co-occur at different stages of protein folding.
Broadly, these modifications can be categorized into four main categories — glycosylation, formation of disulfide bonds, assembly of protein subunits, and specific proteolytic cleavages like removal of signal...
Broadly, these modifications can be categorized into four main categories — glycosylation, formation of disulfide bonds, assembly of protein subunits, and specific proteolytic cleavages like removal of signal...
6.8K


