Related Experiment Videos
PPO-GPR: A Custom Proximal Policy Optimization Tool for Active Reinforcement Learning
Etinosa Osaro1, Yamil J Colón1
1Department of Chemical and Biomolecular Engineering, University of Notre Dame, Notre Dame, Indiana 46556, United States.
Abstract:
Efficient data selection is critical in domains where data acquisition is expensive and time-consuming, such as material science. In this work, we introduce a novel active learning framework that integrates proximal policy optimization (PPO) with Gaussian process regression (GPR) to strategically select informative data points and thereby enhance predictive modeling. Leveraging the inherent stability and sample efficiency of PPO, achieved through a clipped surrogate objective, the framework guides data acquisition via a custom-designed Gymnasium environment tailored for GPR. In this environment, the PPO agent dynamically chooses data points based on their potential to improve the GPR's performance, as measured by the R 2 score, while preventing redundancy through an action masking mechanism. We apply the proposed methodology to predict the selectivity of methane (CH4) over higher alkanes in metal-organic frameworks (MOFs), focusing on CuBTC and IRMOF-1. The framework is evaluated using both ternary and quaternary gas mixtures, where the performance of the GPR is assessed through metrics such as R 2, mean absolute error (MAE), and root mean squared error (RMSE). Across CuBTC and IRMOF-1 in ternary and quaternary hydrocarbon mixtures, PPO-guided acquisition achieves 77-86% data savings relative to full GCMC grids, typically querying only ∼14-23% of the candidate pool while the clipped-update PPO policy converges stably by focusing selections in the pressure-temperature-composition regions where selectivity changes most rapidly. This work shows the potential of combining advanced reinforcement learning techniques with regression models to accelerate material discovery and optimize gas separation processes.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can have a...
Reinforcement Schedules
Once a behavior is learned,...
Observational Learning
Randomized Experiments
Simple randomization
Simple...
Rolling Resistance: Problem Solving