Related Experiment Video
Updated: May 24, 2025

08:18
WheelCon: A Wheel Control-Based Gaming Platform for Studying Human Sensorimotor Control
Published on: August 15, 2020
4.9K
Inverse Reinforcement Learning for Discrete-Time Systems With Data Dropouts
IEEE Transactions on Cybernetics
|March 4, 2025
Summary
This study introduces inverse reinforcement learning (IRL) algorithms to enable control systems to track target systems effectively, even with data loss during wireless transmission. The methods allow systems to learn unknown target behaviors for improved tracking performance.
Area of Science:
- Control Systems Engineering
- Machine Learning
- Wireless Communication
Background:
- Networked control systems (NCS) face challenges with data loss during wireless transmission.
- Tracking control requires understanding an unknown target system's behavior, often defined by an unknown cost function.
Purpose of the Study:
- To develop inverse reinforcement learning (IRL) algorithms for tracking control of linear NCS with random state dropouts.
- To enable a controlled system to infer the unknown cost function and optimal policy of a target system.
Main Methods:
- A model-based IRL algorithm integrating a Smith predictor for state estimation was developed.
- A state-dropout-aware inverse Q-learning algorithm was proposed, requiring only accessible system data.
Main Results:
- The proposed algorithms effectively infer the target's cost function and optimal control policy.
- Theoretical validity was rigorously established.
- Numerical simulations confirmed the practical effectiveness of the algorithms.
Conclusions:
- The developed IRL algorithms provide a robust solution for tracking control in NCS with random state dropouts.
- These methods enhance tracking performance by enabling systems to learn and adapt to unknown target dynamics despite communication uncertainties.
Related Concept Videos
Reinforcement Schedules
126
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
126
Classification of Systems-II
133
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
133
Second Order systems II
79
In an underdamped second-order system, where the damping ratio ζ is between 0 and 1, a unit-step input results in a transfer function that, when transformed using the inverse Laplace method, reveals the output response. The output exhibits a damped sinusoidal oscillation, and the difference between the input and output is termed the error signal. This error signal also demonstrates damped oscillatory behavior. Eventually, as the system reaches a steady state, the error diminishes to zero.
79
Sampling Continuous Time Signal
195
In signal processing, a continuous-time signal can be sampled using an impulse-train sampling technique, followed by the zero-order hold method. Impulse-train sampling involves the use of a periodic impulse train, which consists of a series of delta functions spaced at regular intervals determined by the sampling period. When a continuous-time signal is multiplied by this impulse train, it generates impulses with amplitudes corresponding to the signal's values at the sampling points.
In the...
In the...
195
Feedback control systems
268
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
268
BIBO stability of continuous and discrete -time systems
321
System stability is a fundamental concept in signal processing, often assessed using convolution. For a system to be considered bounded-input bounded-output (BIBO) stable, any bounded input signal must produce a bounded output signal. A bounded input signal is one where the modulus does not exceed a certain constant at any point in time.
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
To determine the BIBO stability, the convolution integral is utilized when a bounded continuous-time input is applied to a Linear Time-Invariant (LTI) system....
321

