Related Experiment Video
Updated: Jul 15, 2026

Protein WISDOM: A Workbench for In silico De novo Design of BioMolecules
Published on: July 25, 2013
Understanding language model scaling for protein fitness prediction
Chao Hou1, Di Liu2, Aziz Zafar2
1Department of Systems Biology, Columbia University Irving Medical Center, New York, NY, USA. ch3849@cumc.columbia.edu.
Protein language models' fitness prediction performance decreases with increasing size. Larger models may overestimate sequence likelihood, hindering accurate mutation effect prediction and protein design.
Area of Science:
- Computational biology
- Deep learning for proteins
- Bioinformatics
Background:
- Protein language models (PLMs) estimate sequence likelihood (p(sequence)) for fitness prediction.
- Larger deep learning models are generally assumed to perform better.
- PLM performance in fitness prediction declines beyond a certain model size.
Purpose of the Study:
- Investigate the impact of model size on protein fitness prediction.
- Clarify the scalability of PLMs for predicting mutation effects.
- Provide guidelines for optimizing PLM application in protein design.
Main Methods:
- Analyzed the relationship between PLM size and predicted sequence likelihood (p(sequence)).
- Evaluated how model size, training data, and stochasticity affect fitness landscape representation.
- Compared predicted p(sequence) with evolutionary patterns in homologous proteins.
Main Results:
- Model size, training data, and stochasticity can bias predicted p(sequence) from true protein fitness.
- Optimal fitness prediction occurs when predicted p(sequence) moderately matches evolutionary patterns.
- Larger models tend to predict excessively high p(sequence), reducing fitness prediction accuracy.
- Extreme predicted likelihoods lead to uniform predictions for mutations, failing to capture the real fitness landscape.
Conclusions:
- Protein language model performance in fitness prediction is not solely dependent on size.
- Moderate predicted sequence likelihoods are crucial for accurately reflecting protein fitness landscapes.
- Findings offer practical guidance for developing and applying protein models in computational biology and protein engineering.
Related Concept Videos
Physiological Pharmacokinetic Models: Assumption with Protein Binding
Improving Translational Accuracy
Improving Translational Accuracy
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Scaling
Mechanistic Models: Compartment Models in Individual and Population Analysis

