Related Experiment Video
Updated: Jun 29, 2025

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
Neural Q-learning for discrete-time nonlinear zero-sum games with adjustable convergence rate.
Yuan Wang1, Ding Wang1, Mingming Zhao1
1Faculty of Information Technology, Beijing University of Technology, Beijing 100124, China; Beijing Key Laboratory of Computational Intelligence and Intelligent System, Beijing University of Technology, Beijing 100124, China; Beijing Institute of Artificial Intelligence, Beijing University of Technology, Beijing 100124, China; Beijing Laboratory of Smart Environmental Protection, Beijing University of Technology, Beijing 100124, China.
This study introduces an adjustable Q-learning scheme to speed up solving nonlinear zero-sum games using neural networks. The developed algorithms ensure convergence and demonstrate excellent performance in simulations.
Area of Science:
- Control Theory
- Artificial Intelligence
- Game Theory
Background:
- Zero-sum games present complex control challenges.
- Model-free approaches are needed for practical applications.
- Accelerating iterative learning algorithms is crucial for efficiency.
Purpose of the Study:
- To develop an adjustable Q-learning scheme for discrete-time nonlinear zero-sum games.
- To enhance the convergence rate of Q-function iteration.
- To enable model-free tracking control using neural networks.
Main Methods:
- Analysis of monotonicity and convergence for iterative Q-functions.
- Integration of neural networks for model-free control.
- Design of two accelerated Q-learning algorithms with convergence guarantees.
Main Results:
- An adjustable Q-learning scheme accelerates convergence.
- Two algorithms ensure convergence with adaptive acceleration phases or adjustable relaxation factors.
- Simulations confirm the algorithm's effectiveness in a practical scenario.
Conclusions:
- The proposed adjustable Q-learning scheme effectively solves nonlinear zero-sum games.
- Neural network integration facilitates model-free tracking control.
- The developed algorithms offer accelerated and guaranteed convergence for Q-learning.
More Related Videos
Related Concept Videos
Current Growth And Decay In RL Circuits
BIBO stability of continuous and discrete -time systems
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
Feedback control systems
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear time-invariant Systems
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
Linear Approximation in Time Domain
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
Region of Convergence

