Related Experiment Video
Updated: Jun 21, 2025

Synthesis and Performance Characterizations of Transition Metal Single Atom Catalyst for Electrochemical CO2 Reduction
Published on: April 10, 2018
Optimal Dynamic Regimes for CO Oxidation Discovered by Reinforcement Learning.
Mikhail S Lifar1, Andrei A Tereshchenko1, Aleksei N Bulgakov1
1The Smart Materials Research Institute, Southern Federal University, 344090 Rostov-on-Don, Russia.
Reinforcement learning optimized CO oxidation on palladium catalysts by dynamically adjusting conditions. This approach uncovered optimal stationary, periodic, and nonperiodic reaction regimes for improved CO2 production.
Area of Science:
- Catalysis
- Chemical Engineering
- Artificial Intelligence
Background:
- Metal nanoparticles are crucial heterogeneous catalysts for activating molecules and lowering reaction energy barriers.
- Reaction yield is influenced by adsorption, activation, desorption, and reaction dynamics, which depend on gas composition, temperature, and pressure.
- Steady-state conditions can lead to catalyst deactivation; dynamic control offers potential for improved performance.
Purpose of the Study:
- To apply reinforcement learning (RL) for dynamic control of CO oxidation on a palladium catalyst.
- To investigate optimal control strategies, including stationary, periodic, and nonperiodic regimes.
- To demonstrate the benefits of dynamic control in heterogeneous catalysis.
Main Methods:
- Utilized a policy gradient reinforcement learning algorithm trained in a theoretical environment.
- Parametrized the RL model using experimental data for CO oxidation.
- Controlled the reaction based on CO and O2 partial pressures over successive time steps.
Main Results:
- The RL algorithm successfully learned to maximize the CO2 formation rate.
- Identified optimal stationary, periodic, and nonperiodic control regimes.
- Provided insights into the advantages of dynamic control for catalytic processes.
Conclusions:
- Reinforcement learning is a viable approach for optimizing heterogeneous catalytic reactions.
- Dynamic control strategies, particularly periodic and nonperiodic regimes, can outperform stationary conditions.
- This work promotes the adoption of RL in catalytic science for enhanced process control and efficiency.
More Related Videos
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Redox Equilibria: Overview
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Ladder Diagrams: Redox Equilibria
Consider the Fe3+/Fe2+ half-reaction, which has a standard-state potential of +0.771 V. At potentials more positive than +0.771 V, Fe3+ predominates, whereas Fe2+...
Dynamic Equilibrium
Observational Learning

