Related Experiment Video
Updated: Jul 19, 2025

A Protocol for Functional Assessment of Whole-Protein Saturation Mutagenesis Libraries Utilizing High-Throughput Sequencing
Published on: July 3, 2016
Learning protein fitness landscapes with deep mutational scanning data from multiple sources
Lin Chen1, Zehong Zhang1, Zhenghao Li2
1Drug Discovery and Design Center, State Key Laboratory of Drug Research, Shanghai Institute of Materia Medica, Chinese Academy of Sciences, Shanghai 201203, China; University of Chinese Academy of Sciences, Beijing 100049, China.
This study introduces a multi-protein training scheme to improve machine learning-assisted directed evolution (MLDE) by leveraging existing data to map protein fitness landscapes more accurately. The findings highlight potential pitfalls in MLDE and suggest better approaches for protein engineering.
Area of Science:
- Biochemistry
- Computational Biology
- Machine Learning
Background:
- Accurate fitness landscape mapping is crucial for machine learning-assisted directed evolution (MLDE).
- Existing MLDE methods often struggle with generalizing to new protein targets.
- Deep mutational scanning (DMS) provides valuable data for understanding protein sequence-function relationships.
Purpose of the Study:
- To develop and validate a multi-protein training scheme for improving fitness landscape prediction in MLDE.
- To assess the scheme's ability to generalize to new proteins and predict higher-order variant effects.
- To identify potential limitations and pitfalls in current MLDE approaches.
Main Methods:
- A novel multi-protein training scheme was designed, utilizing existing DMS data from diverse proteins.
- Proof-of-concept trials were conducted to validate the scheme through random and positional extrapolation, zero-shot predictions, and higher-order effect extrapolation.
- Performance was benchmarked against strong baseline models.
Main Results:
- The multi-protein training scheme demonstrated effectiveness in aiding the understanding of new protein fitness landscapes.
- The study validated the scheme's performance in random/positional extrapolation, zero-shot predictions, and higher-order effect prediction.
- Unexpectedly strong performance from baseline models revealed significant pitfalls in current MLDE practices.
Conclusions:
- The proposed multi-protein training scheme offers a promising approach to enhance MLDE by improving fitness landscape learning.
- The findings underscore the importance of robust baselines and highlight areas for improvement in MLDE methodologies.
- This work contributes to a better understanding of protein fitness profiles and advances the development of more effective protein engineering tools.

