Related Experiment Video
Updated: Jun 14, 2025

WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
Asynchronous iterative Q-learning based tracking control for nonlinear discrete-time multi-agent systems
Ziwen Shen1, Tao Dong1, Tingwen Huang2
1College of Electronics and Information Engineering, Southwest University, Chongqing, 400715, PR China.
This study introduces a new asynchronous iterative Q-learning (AIQL) algorithm for tracking control in nonlinear discrete-time multi-agent systems (MASs). The AIQL algorithm demonstrates faster convergence and lower costs compared to existing methods.
Area of Science:
- Control Systems Engineering
- Artificial Intelligence
- Robotics
Background:
- Multi-agent systems (MASs) present complex challenges in coordinated control.
- Tracking control is crucial for MASs to follow desired trajectories.
- Existing algorithms may lack efficiency in nonlinear discrete-time MASs.
Purpose of the Study:
- To develop a novel tracking control algorithm for nonlinear discrete-time MASs.
- To transform the tracking problem into an optimal regulation problem of a local neighborhood error system (LNES).
- To implement the algorithm using a neural network-based actor-critic framework.
Main Methods:
- Construction of a local neighborhood error system (LNES).
- Development of an asynchronous iterative Q-learning (AIQL) algorithm with two Q-values (QiA and QiB) for policy improvement and evaluation.
- Implementation of AIQL using a neural network-based actor-critic framework with two critic networks.
Main Results:
- The LNES is shown to converge to 0, effectively solving the tracking problem.
- The AIQL-based algorithm achieved a lower cost value compared to the iterative Q-learning (IQL) based algorithm.
- The AIQL algorithm exhibited a faster convergence speed than the IQL-based algorithm.
Conclusions:
- The proposed AIQL algorithm provides an effective solution for tracking control in nonlinear discrete-time MASs.
- The AIQL algorithm offers improved performance in terms of cost and convergence speed.
- The neural network-based actor-critic implementation successfully demonstrates the algorithm's capabilities.
Related Concept Videos
Open and closed-loop control systems
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
Multi-input and Multi-variable systems
In the absence...
Feedback control systems
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear time-invariant Systems
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
One-Degree-of-Freedom System
A one-degree-of-freedom system is defined by an independent variable that determines its state and behavior. One example of a one-degree-of-freedom system is a simple harmonic oscillator, such as a...
BIBO stability of continuous and discrete -time systems
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....

