Related Experiment Video
Updated: Apr 6, 2026

A Modified Lean and Release Technique to Emphasize Response Inhibition and Action Selection in Reactive Balance
Published on: March 19, 2020
Offline constrained policy optimization with safe anchoring
Diyuan Hou1, Longyang Huang2, Pu Feng1
1School of Computer Science and Engineering, Beihang University, China.
Abstract:
The application of reinforcement learning (RL) in real-world scenarios is limited due to safety concerns and the distribution-shift challenge in the offline setting. To address these issues, we formulate safe offline RL as a constrained policy optimization problem that integrates cumulative cost constraints and behavioral policy regularization. We first derive the analytical solution of the offline constrained policy optimization problem through Lagrangian duality. Then, we prove that iterative updates of this solution guarantee monotonic performance improvement while bounding worst-case costs relative to the behavioral policy. To further prevent out-of-distribution actions that may violate safety constraints, we propose a mechanism that distills a "safe action" distribution from the offline data and restricts policy updates within this safe region. We term this approach safe anchoring. By projecting the analytical solution into a parameterized policy space using a VAE-distilled safe anchoring mechanism, we develop the Offline Constrained Policy Optimization with Safe Anchoring (OCPO-SA) algorithm. Extensive experiments on Safety-Gymnasium and Bullet-Safety-Gym demonstrate that OCPO-SA achieves safety in all tested environments, with the average cost reduced by 24% compared with the best-performing baseline among the compared algorithms.
Related Concept Videos
The Anchoring-and-Adjustment Heuristic
Anchoring Junctions
Constraints and Statical Determinacy
Stability of Equilibrium Configuration: Problem Solving
Problem-solving in the context of the stability of equilibrium configuration...
Pole and System Stability
Simple poles are unique roots of the denominator polynomial. Each simple pole corresponds to a distinct solution to the system's characteristic equation, typically resulting in exponential decay terms in the system's...
Statically Indeterminate Problem Solving