Related Experiment Video
Updated: May 14, 2026

Generic Protocol for Optimization of Heterologous Protein Production Using Automated Microbioreactor Technology
Published on: December 15, 2017
Dynamic, Unconstrained Optimization of Secreted Enzyme Production in Fed-Batch Fermentation Using Reinforcement
Sai Harish Uthravalli1, Sakib Ferdous2, J Michael Hess3
1Department of Computer Science, College of Liberal Arts and Sciences, Iowa State University, Ames, Iowa, USA.
None:
Reinforcement learning (RL) has been used to control a wide range of dynamic processes, especially ones that are too complex to model well or have stochastic environmental perturbations. Fed-batch fermentations are subject to changes in starting cell growth rates and process variations that can affect cell growth and secreted target production. RL has been shown on digital environments of fermentation to control known setpoints (such as temperature) but has yet to be demonstrated for unconstrained product maximization. In this work we develop a fed-batch fermentation model (digital twin) of Aspergillus niger secreting glucoamylase using the Monod model, known literature parameters, and assumed constants to align with typical production values. An RL agent is trained on this environment to evaluate types of algorithms (Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC)), rate of learning, and effects of process perturbations. State variables fed to the model include run time, cell concentration, and measured enzyme activity in the fermentation broth, with the objective of maximizing the enzyme activity. It is found that SAC outperforms PPO, achieving 77.7% of the maximum quality with 200 training episodes and 95% at the 2400th episode, compared to PPO which achieves 80% of the max reward after 2912 episodes of training. The RL controller is benchmarked against a traditional, model-free controller that used Bayesian optimization to discover the optimal feed rate for a given cell type. The traditional controller can be implemented with fewer training runs; however, it is not as robust when exposed to variations in starting cell growth or process perturbations including faulty feed or cooling pumps. In all cases, the RL controller can maintain higher enzyme production, despite changes in the process. Finally, the RL controller is exposed to new cell types (in silico) to determine the experimental cost of updating the trained model with real bioreactor runs. Surprisingly, we found that with no updates the model can perform well across a wide range of new cell types, and that by retraining the quality of performance improves. These results indicate that an in silico trained RL agent can be updated with an array of fermentation experiments to provide robust fermentation control.
Related Concept Videos
Bioreactor Controls-III
Upstream Processing
Batch vs Continuous Culture
Fed-Batch Culture
Enzyme Kinetics
Scientists typically study enzyme kinetics with a fixed amount of enzyme in the controlled environment of a test tube. When more reactant, or substrate, is...
Designing Growth Media for Bioreactors

