LSTM-empowered reinforcement learning in Bi-level optimal control for nonlinear systems with uncertain dynamics
Roya Khalili Amirabadi1, Mohsen Jalaeian-Farimani2, Omid S Fard1
1Department of Applied Mathematics, Ferdowsi University of Mashhad, Mashhad, Iran.
ISA Transactions
|November 27, 2025
Summary
This study presents a novel bi-level optimization framework for robust control of nonlinear systems. It uses Long Short-Term Memory networks and reinforcement learning for adaptive trajectory tracking in uncertain environments.
Area of Science:
- Robotics and Control Systems
- Artificial Intelligence
- Nonlinear System Dynamics
Background:
- Autonomous systems require robust control strategies for reliable operation in dynamic and uncertain environments.
- Traditional control methods often struggle with time-varying disturbances and complex nonlinear dynamics.
- Reinforcement learning (RL) and deep learning offer promising avenues for adaptive control but require careful integration and stability guarantees.
Purpose of the Study:
- To develop a novel bi-level optimization framework for optimal control of nonlinear continuous-time systems with uncertain dynamics.
- To integrate Long Short-Term Memory (LSTM) networks with an actor-critic reinforcement learning (RL) architecture for enhanced adaptive control.
- To achieve robust trajectory tracking without the need for offline training, ensuring adaptability to time-varying disturbances.
Main Methods:
- A bi-level optimization framework combining Hamiltonian-based optimal control with online uncertainty estimation.
- Utilizing an actor-critic RL architecture where the master level optimizes control policies (HJB-inspired) and the slave level uses LSTM networks for dynamic uncertainty estimation.
- Implementing rigorous stability analysis to prove uniform ultimate boundedness of the tracking error.
Main Results:
- Demonstrated superior tracking precision, energy efficiency, and disturbance rejection capabilities compared to conventional adaptive control and model-based virtual reference trajectory schemes.
- Validated the framework's effectiveness through extensive simulations on a skid-steering tracked robot executing diverse trajectories.
- Confirmed the framework's ability to handle time-varying disturbances and adapt to uncertain system dynamics.
Conclusions:
- The proposed bi-level optimization framework offers a computationally efficient and theoretically grounded solution for optimal control of nonlinear systems with uncertainties.
- This approach advances the paradigm of RL-based optimal control, providing a scalable solution for autonomous systems operating in unpredictable environments.
- The integration of LSTM networks and actor-critic RL ensures robust and adaptive trajectory tracking, enhancing system performance and reliability.
Related Concept Videos
Linear time-invariant Systems
846
A system is linear if it displays the characteristics of homogeneity and additivity, together termed the superposition property. This principle is fundamental in all linear systems. Linear time-invariant (LTI) systems include systems with linear elements and constant parameters.
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
The input-output behavior of an LTI system can be fully defined by its response to an impulsive excitation at its input. Once this impulse response is known, the system's reaction to any other input can be...
846
Linear Approximation in Time Domain
323
Nonlinear systems often require sophisticated approaches for accurate modeling and analysis, with state-space representation being particularly effective. This method is especially useful for systems where variables and parameters vary with time or operating conditions, such as in a simple pendulum or a translational mechanical system with nonlinear springs.
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
For a simple pendulum with a mass evenly distributed along its length and the center of mass located at half the pendulum's length,...
323
Feedback control systems
663
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
663
Open and closed-loop control systems
1.5K
Control systems are foundational elements in automation and engineering. They are broadly categorized into open-loop and closed-loop systems. These classifications hinge on the presence or absence of feedback mechanisms, significantly influencing the system's performance, complexity, and application.
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
1.5K
State Space Representation
502
The frequency-domain technique, commonly used in analyzing and designing feedback control systems, is effective for linear, time-invariant systems. However, it falls short when dealing with nonlinear, time-varying, and multiple-input multiple-output systems. The time-domain or state-space approach addresses these limitations by utilizing state variables to construct simultaneous, first-order differential equations, known as state equations, for an nth-order system.
Consider an RLC circuit, a...
Consider an RLC circuit, a...
502
Control Systems
1.8K
Control systems are everywhere in contemporary society, influencing diverse applications from aerospace to automated manufacturing. These systems can be found naturally within biological processes, such as blood sugar regulation and heart rate adjustment in response to stress, as well as in man-made systems like elevators and automated vehicles. A control system is essentially a network of subsystems and processes that collaboratively convert specific inputs into desired outputs.
At the heart...
At the heart...
1.8K


