Related Experiment Video
Updated: Jul 2, 2025

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
5.0K
Model-Free Q-Learning for the Tracking Problem of Linear Discrete-Time Systems
IEEE Transactions on Neural Networks and Learning Systems
|February 21, 2024
Summary
A novel model-free Q-learning algorithm addresses tracking errors in linear systems with unknown dynamics. This approach transforms tracking into regulation, ensuring accurate control policy convergence with minimal data.
Area of Science:
- Control Systems Engineering
- Machine Learning
- Robotics
Background:
- Tracking problems in linear discrete-time systems often suffer from unknown system dynamics, hindering precise control.
- Existing adaptive dynamic programming (ADP) and Q-learning methods have limitations in handling completely unknown system parameters for tracking tasks.
Purpose of the Study:
- To propose a model-free Q-learning algorithm for solving the tracking problem in linear discrete-time systems with unknown dynamics.
- To develop a novel performance index that transforms the tracking problem into a regulation problem, thereby eliminating tracking errors.
Main Methods:
- A model-free Q-learning algorithm is introduced, utilizing an enhanced performance index with an added product term.
- The control policy is deduced iteratively using online system state, control input, and reference trajectory information, without prior system knowledge.
- An off-policy approach is incorporated to optimize data usage for deriving the optimal control policy.
Main Results:
- The proposed performance index effectively converts the tracking problem into a regulation problem, aiming for zero steady-state error.
- The iterative Q-learning method, combined with the new performance index, yields a control policy that theoretically eliminates tracking errors.
- Numerical simulations validate the effectiveness of the proposed algorithm in achieving accurate system tracking.
Conclusions:
- The developed model-free Q-learning algorithm offers a robust solution for tracking control in systems with unknown dynamics.
- The novel performance index and iterative update strategy ensure accurate tracking and efficient policy learning.
- This approach provides a significant advancement over existing methods, particularly in scenarios with limited system information.
Related Concept Videos
Linear Approximation in Time Domain
81
Nonlinear systems often require sophisticated approaches for accurate modeling and analysis, with state-space representation being particularly effective. This method is especially useful for systems where variables and parameters vary with time or operating conditions, such as in a simple pendulum or a translational mechanical system with nonlinear springs.
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
81
Linear time-invariant Systems
258
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
258
Feedback control systems
313
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
313
Linear Approximation in Frequency Domain
91
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
91
First Order Systems
91
First-order systems, such as RC circuits, are foundational in understanding dynamic systems due to their straightforward input-output relationship. Analyzing their responses to different input functions under zero initial conditions reveals significant insights into system behavior.
When a first-order system is subjected to a unit-step input, its response is characterized by its transfer function. By applying the Laplace transform of the unit-step input to the transfer function, expanding the...
When a first-order system is subjected to a unit-step input, its response is characterized by its transfer function. By applying the Laplace transform of the unit-step input to the transfer function, expanding the...
91
State Space Representation
208
The frequency-domain technique, commonly used in analyzing and designing feedback control systems, is effective for linear, time-invariant systems. However, it falls short when dealing with nonlinear, time-varying, and multiple-input multiple-output systems. The time-domain or state-space approach addresses these limitations by utilizing state variables to construct simultaneous, first-order differential equations, known as state equations, for an nth-order system.
Consider an RLC circuit, a...
Consider an RLC circuit, a...
208

