Related Experiment Video
Updated: Mar 27, 2026

Real-Time Proxy-Control of Re-Parameterized Peripheral Signals using a Close-Loop Interface
Published on: May 8, 2021
Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for safe reinforcement learning in
Roya Khalili-Amirabadi1, Mohsen Jalaeian-Farimani2, Omid Solaymani-Fard3
1Department of Applied Mathematics, Ferdowsi University of Mashhad, Mashhad, Iran. roya.khalili.a@gmail.com.
This study introduces Self-Organizing Dual-buffer Adaptive Clustering Experience Replay (SODACER) for safe optimal control. SODACER enhances learning efficiency and safety in complex systems.
Area of Science:
- Control Engineering
- Machine Learning
- Artificial Intelligence
Background:
- Optimal control of nonlinear systems presents challenges in safety and scalability.
- Reinforcement learning (RL) often struggles with efficient memory management and maintaining diverse experiences.
Purpose of the Study:
- To propose a novel RL framework, SODACER, for safe and scalable optimal control of nonlinear systems.
- To integrate experience replay mechanisms with safety constraints for robust learning.
Main Methods:
- Developed Self-Organizing Dual-buffer Adaptive Clustering Experience Replay (SODACER) with Fast and Slow Buffers.
- Implemented a self-organizing adaptive clustering mechanism for experience replay memory optimization.
- Integrated SODACER with Control Barrier Functions (CBFs) for safety guarantees and the Sophia optimizer for enhanced convergence.
Main Results:
- SODACER demonstrated faster convergence and improved sample efficiency compared to baseline methods.
- The framework achieved a superior bias-variance trade-off while ensuring safe system trajectories.
- Validated effectiveness on a nonlinear Human Papillomavirus (HPV) transmission model with safety constraints.
Conclusions:
- SODACER provides a reliable, effective, and robust learning architecture for dynamic, safety-critical environments.
- The approach offers a generalizable solution for robotics, healthcare, and large-scale system optimization.
- The proposed method significantly enhances learning performance and safety in complex control tasks.
Related Concept Videos
Observational Learning
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Associative Learning
Classical conditioning, also known...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Reinforcement Schedules
Once a behavior is learned,...
Multi-input and Multi-variable systems
In the absence of...
