Simulated Design-Build-Test-Learn Cycles for Consistent Comparison of Machine Learning Methods in Metabolic
Paul van Lent1, Joep Schmitz2, Thomas Abeel1,3
1Delft Bioinformatics Lab, Delft University of Technology Van Mourik, Delft 2628 XE, The Netherlands.
ACS Synthetic Biology
|August 24, 2023
Summary
Machine learning aids metabolic flux optimization via iterative design-build-test-learn cycles. A new framework shows gradient boosting and random forest models excel in low-data scenarios, improving strain development.
Area of Science:
- Metabolic Engineering and Synthetic Biology
- Computational Biology and Bioinformatics
- Machine Learning in Biotechnology
Background:
- Combinatorial pathway optimization is crucial for metabolic flux optimization but faces combinatorial explosions with many genes.
- Iterative design-build-test-learn (DBTL) cycles are commonly used for strain optimization, aiming for incremental product strain development.
- Evaluating machine learning (ML) methods for iterative DBTL cycles is challenging due to a lack of standardized testing frameworks.
Purpose of the Study:
- To propose and validate a mechanistic kinetic model-based framework for testing and optimizing ML in iterative combinatorial pathway optimization.
- To assess the performance of various ML models in the context of DBTL cycles for strain optimization.
- To develop an algorithm for recommending new designs based on ML predictions within limited strain construction budgets.
Main Methods:
- Developed a mechanistic kinetic model-based framework to simulate and evaluate ML performance across multiple DBTL cycles.
- Tested and compared various ML algorithms, including gradient boosting and random forest, under low-data conditions.
- Introduced a design recommendation algorithm leveraging ML model predictions for efficient strain selection.
Main Results:
- Gradient boosting and random forest models demonstrated superior performance compared to other ML methods in low-data regimes.
- These robust ML models proved resilient to training set biases and experimental noise inherent in biological data.
- The study found that initiating with a larger initial DBTL cycle is more advantageous when the number of constructible strains is limited.
Conclusions:
- The proposed framework provides a robust method for evaluating and optimizing ML for iterative strain development.
- Gradient boosting and random forest are effective ML choices for low-data, noisy environments in metabolic engineering.
- Strategic DBTL cycle design, particularly larger initial cycles, can enhance efficiency in strain optimization efforts.


