Accurate top protein variant discovery via low-N pick-and-validate machine learning
Hoi Yee Chu1, John H C Fong2, Dawn G L Thean2
1Laboratory of Combinatorial Genetics and Synthetic Biology, School of Biomedical Sciences, The University of Hong Kong, Pokfulam, Hong Kong SAR, China; Centre for Oncology and Immunology, Hong Kong Science Park, Hong Kong SAR, China.
Cell Systems
|February 10, 2024
Summary
This study introduces a machine learning strategy for protein engineering, efficiently identifying top-performing variants with minimal experiments. The method uses active learning to boost resource production and protein design.
Area of Science:
- Protein Engineering
- Machine Learning
- Computational Biology
Background:
- Optimizing protein variants is crucial for biotechnology.
- Vast combinatorial landscapes pose experimental challenges.
- Efficiently identifying high-performing mutants requires advanced strategies.
Purpose of the Study:
- To develop a machine learning-based strategy for efficient protein variant selection.
- To minimize experimental effort in identifying best-performing protein variants.
- To enhance resource producibility in protein engineering.
Main Methods:
- Integrating zero-shot prediction with multi-round active learning.
- Employing low-N pick-and-validate sampling for experimental design.
- Utilizing machine learning to predict and select top variants.
Main Results:
- Achieved up to 92.6% accuracy in selecting the top 1% of variants.
- Demonstrated effectiveness with four rounds of 12-variant sampling.
- Showed comparable results with two rounds of 24-variant sampling.
Conclusions:
- The proposed strategy outperforms existing methods in protein variant discovery.
- Successfully identified high-performance variants in CRISPR-based genome editors.
- The approach is generalizable for various protein engineering tasks.


