Related Experiment Video
Updated: Sep 3, 2025

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
5.0K
Nearly Optimal Control for Mixed Zero-Sum Game Based on Off-Policy Integral Reinforcement Learning
Summary
This study introduces an integral reinforcement learning (IRL) algorithm for mixed zero-sum games with unknown nonlinear system dynamics. The method finds optimal control strategies for competitors and collaborators without needing system information.
Area of Science:
- Control Theory
- Game Theory
- Machine Learning
Background:
- Solving mixed zero-sum games with unknown nonlinear system dynamics presents significant challenges.
- Traditional methods often require complete system information, limiting their applicability.
Purpose of the Study:
- To develop a novel policy iterative algorithm using integral reinforcement learning (IRL) for mixed zero-sum games.
- To achieve optimal control for competing and collaborating players in systems with unknown dynamics.
Main Methods:
- A policy iterative algorithm employing integral reinforcement learning (IRL) is proposed, which bypasses the need for system information.
- An adaptive update law integrating a critic-actor structure with experience replay is introduced.
- Actor functions are designed to approximate optimal control and estimate auxiliary control simultaneously.
Main Results:
- The proposed algorithm successfully obtains optimal control strategies for all players.
- Parameters of the actor-critic structure are updated simultaneously, ensuring efficient learning.
- Uniform ultimate boundedness of parameter errors in polynomial approximation is mathematically proven.
Conclusions:
- The developed IRL-based algorithm effectively solves mixed zero-sum games with unknown nonlinear dynamics.
- The adaptive critic-actor structure with experience replay offers a robust approach to control optimization.
- Simulation results validate the algorithm's effectiveness and practical applicability.
Related Concept Videos
PI Controller: Design
465
Proportional Integral (PI) controllers are a fundamental component in modern control systems, widely used to enhance performance and mitigate steady-state errors. They are particularly effective in applications such as automatic brightness adjustment on smartphones, where they excel at mitigating steady-state errors for step-function inputs. Unlike PD controllers, which require time-varying errors to function optimally, PI controllers leverage their integral component to address residual...
465
Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model
119
Drugs administered through various routes can lead to nonlinear elimination, resulting in complex pharmacokinetic behaviors crucial to understanding efficacious drug dosing.
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...
119
Time and frequency -Domain Interpretation of PI Control
196
Proportional-Integral (PI) controllers are essential in many control systems to improve stability and performance. They are commonly used in everyday devices like thermostats to enhance system damping and reduce steady-state error. When the zero in the controller's transfer function is optimally placed, the system benefits significantly in terms of stability and accuracy.
Acting as a low-pass filter, the PI controller slows the system's response and extends settling times. This requires...
Acting as a low-pass filter, the PI controller slows the system's response and extends settling times. This requires...
196
Reinforcement
319
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
319
Reinforcement Schedules
232
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
232
Operant Conditioning Intervention
110
Operant conditioning serves as a foundational principle in therapeutic interventions aimed at modifying maladaptive behaviors. Central to this approach is the notion that behaviors, both adaptive and maladaptive, are learned through reinforcement. By analyzing the environmental factors that reinforce problematic behaviors, clinicians can design interventions to weaken these reinforcements and replace maladaptive behaviors with healthier alternatives.
In operant conditioning, behaviors that are...
In operant conditioning, behaviors that are...
110

