Related Experiment Video
Updated: Jan 25, 2026

04:48
Control of Eating Behavior Using a Novel Feedback System
Published on: May 8, 2018
11.5K
Output-Feedback Control of Linear Continuous-Time Systems Using Discounted Inverse Reinforcement Learning.
IEEE Transactions on Cybernetics
|January 23, 2026
Summary
This study introduces a new discounted inverse reinforcement learning (DIRL) algorithm for controlling unknown systems using only output data. The method reconstructs states and learns optimal control policies efficiently, outperforming existing techniques.
Area of Science:
- Control Systems Engineering
- Machine Learning
- Robotics
Background:
- Discounted inverse reinforcement learning (DIRL) typically requires full-state feedback, limiting its use in real-world applications with only input-output data.
- Unknown continuous-time (CT) systems with partially observable states present significant control challenges.
- Learning unknown discounted value functions is crucial for optimal control policy derivation.
Purpose of the Study:
- To develop a novel model-free, output-feedback (OPFB) DIRL algorithm for linear quadratic (LQ) control of unknown CT systems.
- To address the limitations of existing DIRL methods by enabling learning from input-output data.
- To reconstruct system states using expert control output data for policy learning.
Main Methods:
- A state reconstruction method is designed utilizing expert control and measured output data.
- A model-free OPFB DIRL algorithm is presented to iteratively learn the unknown value function and optimal control policy.
- Rigorous analysis of algorithm convergence and solution uniqueness is performed.
Main Results:
- The proposed algorithm effectively recovers the expert control policy.
- Simulations demonstrate superior computational efficiency compared to state-of-the-art methods.
- The algorithm successfully handles partially observable states and unknown value functions.
Conclusions:
- The novel OPFB DIRL algorithm provides an effective solution for controlling unknown CT systems with limited state information.
- The method enhances the applicability of DIRL in practical scenarios by utilizing only input-output data.
- The algorithm offers a computationally efficient and robust approach to learning optimal control policies.
Related Concept Videos
Feedback control systems
703
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
703
Linear time-invariant Systems
890
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
890
BIBO stability of continuous and discrete -time systems
898
System stability is a fundamental concept in signal processing, often assessed using convolution. For a system to be considered bounded-input bounded-output (BIBO) stable, any bounded input signal must produce a bounded output signal. A bounded input signal is one where the modulus does not exceed a certain constant at any point in time.
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
898
Linear Momentum in Control Volume
1.3K
Newton's second law is applied to obtain the linear momentum in a control volume in a fluid system. According to this law, the rate of change of linear momentum is equal to the sum of external forces acting on the system. When a control volume matches the fluid system at a specific moment, the forces acting on both are identical. Reynolds transport theorem helps explain this by breaking down the system's linear momentum into two components: the rate of change of linear momentum within...
1.3K
Root Loci for Positive-Feedback Systems
338
The Hartley oscillator is a positive feedback system that sustains oscillations by feeding the output back to the input in phase, thereby reinforcing the signal. Positive feedback systems can be viewed as negative feedback systems with inverted feedback signals. In these systems, the root locus encompasses all points on the s-plane where the angle of the system transfer function equals 360 degrees.
The construction rules for the root locus in positive feedback systems are similar to those in...
The construction rules for the root locus in positive feedback systems are similar to those in...
338
Control Systems
1.8K
Control systems are everywhere in contemporary society, influencing diverse applications from aerospace to automated manufacturing. These systems can be found naturally within biological processes, such as blood sugar regulation and heart rate adjustment in response to stress, as well as in man-made systems like elevators and automated vehicles. A control system is essentially a network of subsystems and processes that collaboratively convert specific inputs into desired outputs.
At the heart...
At the heart...
1.8K

