Related Experiment Video
Updated: Apr 25, 2026

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
Recent Advances on Off-Policy Reinforcement Learning for Optimization Control
Abstract:
Reinforcement learning (RL), a key artificial intelligence technique, has been widely studied and applied over the past two decades to solve various optimization control problems. Generally speaking, there are two basic frameworks for RL-based control design, i.e., on-policy and off-policy RL (OffP-RL). The essential distinction between the two frameworks lies in whether the policy used to generate training data is the behavior policy or the target policy. In on-policy RL-based control methods, the data used for evaluating the target control policy at each iteration must be collected from the system under the target policy itself. In contrast, in OffP-RL methods, the system data is generated by other behavior control policies. It addresses the inadequate exploration problem in on-policy RL methods, making OffP-RL methods more practical and easier to implement. In this article, the recent advances in OffP-RL-based control methods are classified into three categories based on the number of controllers/players involved, i.e., single-/two-/multiplayer. For the single-player case, it is an optimal control problem, which aims to use OffP-RL to learn the optimal control policy, which minimizes the performance index. In the two-player case, most works focus on the $H_{\infty } $ control problem and the two-player zero-sum game, using learning to find the Nash equilibrium. In the multiplayer case, a multiplayer game involves a single system with multiple control inputs, while a multiagent system consists of multiple systems with independent control inputs. Finally, related applications of OffP-RL-based control and future work are analyzed.
Related Concept Videos
Control Systems
At the heart...
Open and closed-loop control systems
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
Reinforcement Schedules
Once a behavior is learned,...
Controller Configurations
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Time-Domain Interpretation of PD Control
Consider the example of control of motor torque. Initially, a positive...