Related Experiment Video
Updated: Dec 27, 2025

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
5.4K
Policy Iteration Q-Learning for Data-Based Two-Player Zero-Sum Game of Linear Discrete-Time Systems
IEEE Transactions on Cybernetics
|February 25, 2020
Summary
A new data-based policy iteration Q-learning (PIQL) algorithm solves zero-sum games for linear systems without needing system dynamics. This method efficiently learns the optimal Q-function using real system data.
Area of Science:
- Control Theory
- Reinforcement Learning
- Game Theory
Background:
- Linear discrete-time systems present challenges in solving two-player zero-sum game problems.
- Traditional methods require complete system dynamics and solving the discrete-time game algebraic Riccati equation (DTGARE).
Purpose of the Study:
- To develop a data-based algorithm for solving zero-sum games in linear discrete-time systems.
- To avoid the need for complete system dynamics and the DTGARE.
Main Methods:
- Introduction of the Q-function to bypass DTGARE.
- Development of a data-based policy iteration Q-learning (PIQL) algorithm using real system data.
- Proof of PIQL algorithm's equivalence to Newton's method using Fréchet derivative and guarantee of convergence via Kantorovich's theorem.
- Implementation using an off-policy learning scheme.
Main Results:
- The PIQL algorithm successfully learns the optimal Q-function from real system data.
- The algorithm's convergence is theoretically guaranteed.
- Simulation studies validate the efficiency of the data-based PIQL method.
Conclusions:
- The data-based PIQL algorithm offers an effective approach to solving zero-sum games for linear discrete-time systems.
- This method eliminates the requirement for system models and DTGARE solutions, enabling practical application.
- The findings demonstrate a significant advancement in data-driven control and learning for game-theoretic problems.
Related Concept Videos
Linear Approximation in Time Domain
284
Nonlinear systems often require sophisticated approaches for accurate modeling and analysis, with state-space representation being particularly effective. This method is especially useful for systems where variables and parameters vary with time or operating conditions, such as in a simple pendulum or a translational mechanical system with nonlinear springs.
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
284
Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model
235
Drugs administered through various routes can lead to nonlinear elimination, resulting in complex pharmacokinetic behaviors crucial to understanding efficacious drug dosing.
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...
235
Linear time-invariant Systems
804
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
804
BIBO stability of continuous and discrete -time systems
843
System stability is a fundamental concept in signal processing, often assessed using convolution. For a system to be considered bounded-input bounded-output (BIBO) stable, any bounded input signal must produce a bounded output signal. A bounded input signal is one where the modulus does not exceed a certain constant at any point in time.
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
843
Feedback control systems
638
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
638
Stability of Equilibrium Configuration: Problem Solving
900
The stability of equilibrium configurations is an important concept in physics, engineering, and other related fields. In simple terms, it refers to the tendency of an object or system to return to its equilibrium position after being disturbed. The stability of an equilibrium configuration can be analyzed by considering the potential energy function of the system and examining its behavior near the equilibrium point.
Problem-solving in the context of the stability of equilibrium configuration...
Problem-solving in the context of the stability of equilibrium configuration...
900