Related Experiment Video
Updated: Jul 12, 2025

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Infer global, predict local: Quantity-relevance trade-off in protein fitness predictions from sequence data
Lorenzo Posani1, Francesca Rizzato1, Rémi Monasson1
1Laboratory of Physics of the Ecole Normale Supérieure, CNRS UMR8023 & PSL Research, Sorbonne Université, Paris, France.
Computational models can predict mutation effects on protein function. Optimal data selection, not just model complexity, is key, revealing that simpler models can excel with well-chosen sequence data.
Area of Science:
- Computational biology
- Protein engineering
- Evolutionary biology
Background:
- Predicting mutation effects on protein function is crucial for evolutionary studies and biomedicine.
- Computational methods, including deep learning, analyze sequence data to predict protein fitness landscapes.
- The interplay between model complexity and data characteristics in prediction accuracy is not well understood.
Purpose of the Study:
- To develop a theoretical framework for understanding prediction error in computational models of protein function.
- To identify key data descriptors that influence predictive performance.
- To determine optimal data selection strategies for improving prediction accuracy.
Main Methods:
- Theoretical analysis of prediction error based on sequence data descriptors.
- Proposing quantity and relevance as key data characteristics.
- Analyzing trade-offs between data quantity, relevance, and model complexity.
- Using repeated subsampling to assess model capture of epistasis.
Main Results:
- A theoretical framework was established to describe prediction error.
- Descriptors for sequence data quantity and relevance were proposed.
- A trade-off between data quantity and relevance was identified, guiding optimal data subset selection.
- Simple models can outperform complex ones when trained on carefully selected data.
- Repeated subsampling reveals the extent to which computational models capture fitness landscape epistasis.
Conclusions:
- Optimal prediction of mutation effects relies on a balance between data quantity and relevance, not solely on model complexity.
- Careful data selection can enhance the performance of simpler computational models.
- The proposed framework provides insights into model limitations regarding epistasis in protein fitness landscapes.
More Related Videos
Related Concept Videos
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conservation of Protein Domains
Protein-protein Interfaces
Gene Evolution - Fast or Slow?
In contrast, regions which code...

