Related Experiment Video
Updated: Dec 9, 2025

08:32
Tracking Rats in Operant Conditioning Chambers Using a Versatile Homemade Video Camera and DeepLabCut
Published on: June 15, 2020
13.1K
Off-Policy Reinforcement Learning for Tracking in Continuous-Time Systems on Two Time Scales
IEEE Transactions on Neural Networks and Learning Systems
|September 9, 2020
Summary
This study uses singular perturbation theory to solve optimal tracking problems in two-time-scale systems. Off-policy integral reinforcement learning effectively approximates solutions for slow and fast subsystems.
Area of Science:
- Control Theory
- Systems Engineering
- Artificial Intelligence
Background:
- Singular perturbation theory is traditionally used for system regulation.
- Optimal linear quadratic tracking (LQT) problems in continuous-time systems present unique challenges.
- Two-time-scale systems require specialized approaches due to differing dynamics.
Purpose of the Study:
- To apply singular perturbation theory to solve the LQT problem for continuous-time two-time-scale processes.
- To demonstrate the separation of the LQT problem into slow and fast subsystem problems.
- To validate the approximation of the original LQT solution by reduced-order control problems.
Main Methods:
- Applying singular perturbation theory to decompose the system.
- Solving the reduced-order linear-quadratic tracker (LQT) problem for the slow subsystem.
- Solving the reduced-order linear-quadratic regulator (LQR) problem for the fast subsystem.
- Utilizing off-policy integral reinforcement learning (IRL) with measured data.
Main Results:
- The two-time-scale tracking problem is successfully separated into independent slow LQT and fast LQR problems.
- The solutions to the reduced-order problems approximate the LQT solution of the original system.
- Off-policy IRL effectively solves these reduced-order control problems using system data.
Conclusions:
- Singular perturbation theory provides an effective framework for solving LQT problems in two-time-scale systems.
- The proposed method, utilizing IRL, offers a data-driven approach to control design.
- Simulation results validate the accuracy and effectiveness of the method compared to traditional approaches.
Related Concept Videos
Basic Continuous Time Signals
582
Basic continuous-time signals include the unit step function, unit impulse function, and unit ramp function, collectively referred to as singularity functions. Singularity functions are characterized by discontinuities or discontinuous derivatives.
The unit step function, denoted u(t), is zero for negative time values and one for positive time values, exhibiting a discontinuity at t=0. This function often represents abrupt changes, such as the step voltage introduced when turning a car's...
The unit step function, denoted u(t), is zero for negative time values and one for positive time values, exhibiting a discontinuity at t=0. This function often represents abrupt changes, such as the step voltage introduced when turning a car's...
582
Sampling Continuous Time Signal
573
In signal processing, a continuous-time signal can be sampled using an impulse-train sampling technique, followed by the zero-order hold method. Impulse-train sampling involves the use of a periodic impulse train, which consists of a series of delta functions spaced at regular intervals determined by the sampling period. When a continuous-time signal is multiplied by this impulse train, it generates impulses with amplitudes corresponding to the signal's values at the sampling points.
In the...
In the...
573
BIBO stability of continuous and discrete -time systems
796
System stability is a fundamental concept in signal processing, often assessed using convolution. For a system to be considered bounded-input bounded-output (BIBO) stable, any bounded input signal must produce a bounded output signal. A bounded input signal is one where the modulus does not exceed a certain constant at any point in time.
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
796
Linear time-invariant Systems
731
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
731
Classification of Systems-II
408
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
408
Reinforcement Schedules
362
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
362

