Fuzzy Lyapunov Reinforcement Learning for Non Linear Systems
Abhishek Kumar1, Rajneesh Sharma1
1Division of Instrumentation and Control Engineering, Netaji Subhas Institute of Technology, Dwarka Sector 3, New Delhi 110078, India.
ISA Transactions
|February 5, 2017
Summary
This study introduces a novel fuzzy reinforcement learning controller that ensures stability by constraining fuzzy rules. It effectively handles complex control problems without needing a system model, demonstrating superior performance.
Area of Science:
- Control Systems Engineering
- Artificial Intelligence
- Fuzzy Logic Systems
Background:
- Traditional control methods often require precise system models and struggle with uncertainties.
- Reinforcement learning (RL) offers adaptive control but ensuring stability, especially in fuzzy systems, remains a challenge.
Purpose of the Study:
- To develop a novel fuzzy reinforcement learning (RL) controller capable of generating stable control actions.
- To introduce Lyapunov constraining of fuzzy rule consequents within an RL framework for guaranteed stability.
- To create a model-free and desired-response-free controller for continuous state-action spaces.
Main Methods:
- Lyapunov constraining the consequent part of fuzzy linguistic rules within a fuzzy RL setup.
- Developing a linguistic RL controller that progressively learns a stable optimal policy.
- Applying the controller to benchmark Inverted Pendulum (IP) and Rotational/Translational Proof-Mass Actuator (RTAC) problems.
Main Results:
- The proposed controller successfully generated stable control actions in continuous state-action spaces.
- Demonstrated effectiveness in handling disturbances on benchmark control problems.
- Simulation results showed superior stability and viability compared to baseline fuzzy Q learning, Lyapunov Actor-Critic, and Lyapunov Markov game controllers.
Conclusions:
- The proposed fuzzy RL controller with Lyapunov constrained consequents is a viable and stable approach for complex control tasks.
- This method offers a robust, model-free solution for systems with uncertainties and disturbances.
- Represents a significant advancement in linguistic RL for achieving stable optimal policies.
Related Concept Videos
Linear Approximation in Frequency Domain
411
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
411
Linear Approximation in Time Domain
387
Nonlinear systems often require sophisticated approaches for accurate modeling and analysis, with state-space representation being particularly effective. This method is especially useful for systems where variables and parameters vary with time or operating conditions, such as in a simple pendulum or a translational mechanical system with nonlinear springs.
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
387
Feedback control systems
755
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
755
Classification of Systems-I
647
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
647
State Space Representation
643
The frequency-domain technique, commonly used in analyzing and designing feedback control systems, is effective for linear, time-invariant systems. However, it falls short when dealing with nonlinear, time-varying, and multiple-input multiple-output systems. The time-domain or state-space approach addresses these limitations by utilizing state variables to construct simultaneous, first-order differential equations, known as state equations, for an nth-order system.
Consider an RLC circuit, a...
Consider an RLC circuit, a...
643
Linear time-invariant Systems
1.0K
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
1.0K

