Related Experiment Video
Updated: Feb 4, 2026

A Method for Remotely Silencing Neural Activity in Rodents During Discrete Phases of Learning
Published on: June 22, 2015
Off-Policy Interleaved Q -Learning: Optimal Control for Affine Nonlinear Discrete-Time Systems
This study introduces a new off-policy interleaved Q-learning algorithm for optimal control of nonlinear systems. The method effectively learns control policies from data without needing system dynamics, ensuring convergence and unbiased solutions.
Area of Science:
- Control Theory
- Machine Learning
- Nonlinear Systems
Background:
- Optimal control of affine nonlinear discrete-time (DT) systems presents significant challenges due to unknown dynamics and the need for off-policy learning.
- Existing on-policy Q-learning methods for these systems have limitations, including potential bias introduced by probing noises required for persistent excitation.
Purpose of the Study:
- To develop a novel off-policy interleaved Q-learning algorithm for optimal control of affine nonlinear DT systems.
- To rigorously prove the convergence of the proposed algorithm and demonstrate its ability to provide unbiased solutions.
Main Methods:
- A review and rigorous convergence proof of on-policy Q-learning for affine nonlinear DT systems.
- Analysis of solution bias in Q-function-based Bellman equations with probing noises.
- Introduction of a behavior control policy and development of an off-policy Q-learning algorithm.
- Implementation of three neural networks within an actor-critic framework using the interleaved Q-learning approach.
Main Results:
- The proposed off-policy interleaved Q-learning algorithm is derived and its convergence is rigorously proven.
- The algorithm effectively addresses challenges posed by unknown dynamics and off-policy learning in nonlinear DT systems.
- Simulation results confirm the effectiveness of the novel algorithm in solving optimal control problems.
Conclusions:
- The novel off-policy interleaved Q-learning algorithm offers a robust solution for optimal control of affine nonlinear discrete-time systems.
- The method successfully overcomes limitations of previous approaches, providing convergence guarantees and unbiased solutions.
- The findings are validated through simulations, demonstrating practical applicability.
More Related Videos
06:04Experimental Investigation of the Hierarchical Control in DC Microgrids Using a Real-time Simulator
Published on: February 14, 2025
08:42Assessment of Social Cognition in Non-human Primates Using a Network of Computerized Automated Learning Device ALDM Test Systems
Published on: May 5, 2015
Related Concept Videos
BIBO stability of continuous and discrete -time systems
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
Discrete-time Fourier transform
One of the notable...
Basic Discrete Time Signals
The unit impulse or sample sequence is mathematically expressed as zero for all n values except at n=0, where it is one. The unit impulse sequence, denoted by δ(n), is the first difference of the unit step sequence, while the unit step sequence u(n) is the...
Discrete-Time Fourier Series
For a discrete-time periodic signal x[n]...
Control Systems
At the heart...
Control Systems: Applications
In modern vehicles, control systems manage various functions to enhance performance and safety. The steering wheel and accelerator are primary inputs in a car's control system. The...