Related Experiment Video
Updated: Jul 1, 2025

09:20
Single-Cell Factor Localization on Chromatin using Ultra-Low Input Cleavage Under Targets and Release using Nuclease
Published on: February 1, 2022
2.7K
Machine learning to predict continuous protein properties from binary cell sorting data and map unseen sequence space
Marshall Case1, Matthew Smith1,2, Jordan Vinh3
1Chemical Engineering, University of Michigan, Ann Arbor, MI 48109.
Summary
This study introduces a machine learning framework to predict protein properties from directed evolution experiments. The approach uses linear models to efficiently identify optimized protein variants with improved functions.
Area of Science:
- Biochemistry and Molecular Biology
- Protein Engineering
- Computational Biology
Background:
- Proteins are essential biomolecules with diverse cellular functions.
- Protein engineering aims to rapidly evolve proteins for improved properties.
- High-throughput methods enhance directed evolution, but data interpretation remains challenging.
Purpose of the Study:
- To develop a framework for predicting continuous protein properties from directed evolution data.
- To improve protein optimization by leveraging interpretable, linear machine learning models.
- To identify lead protein candidates with enhanced functions.
Main Methods:
- Developed a framework using interpretable, linear machine learning models.
- Utilized data from simple, imprecise experimental estimates of protein fitness.
- Applied the framework to predict affinity and specificity from cell sorting data for stapled peptides.
- Coupled integer linear programming with machine learning mutation scores for optimization.
Main Results:
- Linear machine learning models accurately predict continuous protein properties from directed evolution data.
- The framework's predictive capabilities approach those of more precise but expensive methods.
- Protein fitness space is reasonably modeled by linear relationships among sequence mutations.
- Successfully identified protein variants with improved and co-optimal properties in prospective tests.
Conclusions:
- The developed framework offers a versatile tool for analyzing and identifying improved protein variants.
- Predicting continuous protein properties from readily available deep sequencing data is feasible.
- Linear relationships effectively model protein fitness landscapes, enabling efficient optimization.
Related Concept Videos
Signal Sequences and Sorting Receptors
5.4K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.4K
Overview of Protein Sorting and Transport
11.3K
Eukaryotic cells have different membrane-bound organelles with distinct protein requirements. The process by which proteins are targeted to a specific organelle is called protein sorting.
Protein sorting can be of two types: signal-based sorting and vesicle-based trafficking. In signal-based sorting, specific amino acid sequences called sorting signals target proteins to the proper location inside the cell either via gated transport or by protein translocation. In gated transport, folded...
Protein sorting can be of two types: signal-based sorting and vesicle-based trafficking. In signal-based sorting, specific amino acid sequences called sorting signals target proteins to the proper location inside the cell either via gated transport or by protein translocation. In gated transport, folded...
11.3K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K

