Related Experiment Video
Updated: Sep 17, 2025

10:45
A Femtoliter Droplet Array for Massively Parallel Protein Synthesis from Single DNA Molecules
Published on: June 20, 2020
10.4K
AlphaFold distillation for inverse protein design.
Igor Melnyk1, Aurélie Lozano2, Payel Das1
1IBM Research, Yorktown Heights, NY, 10598, USA.
Scientific Reports
|July 2, 2025
Summary
We developed a faster inverse protein folding method using knowledge distillation. This approach improves protein design by creating diverse sequences with consistent structures, aiding bio-engineering and drug discovery.
Area of Science:
- Computational Biology
- Protein Engineering
- Bioinformatics
Background:
- Inverse protein folding designs sequences for specific 3D structures, vital for bio-engineering and drug discovery.
- Existing methods often depend on limited experimental structures.
- Accurate protein structure prediction models (e.g., AlphaFold) are too slow for inverse folding training.
Purpose of the Study:
- To develop a computationally efficient method for inverse protein folding.
- To integrate fast structure prediction into inverse folding model training.
- To enhance sequence recovery and diversity in protein design.
Main Methods:
- Knowledge distillation applied to protein folding model confidence metrics (pTM, pLDDT).
- Creation of a fast, end-to-end differentiable distilled model.
- Utilizing the distilled model as a structure consistency regularizer during inverse folding training.
Main Results:
- The distilled model significantly speeds up structure consistency checks.
- Achieved up to 3% improvement in sequence recovery compared to baselines.
- Demonstrated up to 45% increase in protein diversity while maintaining structural integrity.
Conclusions:
- Knowledge distillation offers an efficient way to regularize inverse folding models.
- The proposed method enhances protein design capabilities for bio-engineering applications.
- This technique is adaptable for other protein design tasks like sequence-based infilling.

