Related Experiment Video
Updated: Feb 6, 2026

Novel Apparatus and Method for Drug Reinforcement
Published on: August 20, 2010
Reinforcement learning via conservative agent for environments with random delays
Jongsoo Lee1, Jangwon Kim1, Jiseok Jeong2
1Department of Convergence IT Engineering, Pohang University of Science and Technology, 77 Cheongam-ro, Nam-gu, Pohang-si, Gyeongbuk, 36763, South Korea.
Abstract:
Real-world reinforcement learning applications are often subject to unavoidable delayed feedback from the environment. Under such conditions, the standard state representation may no longer induce Markovian dynamics unless additional information is incorporated at decision time, which introduces significant challenges for both learning and control. While numerous delay-compensation methods have been proposed for environments with constant delays, those with random delays remain largely unexplored due to their inherent variability and unpredictability. In this study, we propose a robust agent for decision-making under bounded random delays, termed the conservative agent. This agent reformulates the random-delay environment into a constant-delay surrogate, which enables any constant-delay method to be directly extended to random-delay environments without modifying their algorithmic structure. Apart from a maximum delay, the conservative agent does not require prior knowledge of the underlying delay distribution and maintains performance invariant to changes in the delay distribution as long as the maximum delay remains unchanged. We present a theoretical analysis of conservative agent and evaluate its performance on diverse continuous control tasks from the MuJoCo benchmarks. Empirical results demonstrate that it significantly outperforms existing baselines in terms of both asymptotic performance and sample efficiency.
Related Concept Videos
What is Conservation Biology?
Conservation of Small Populations
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Corrosion of Reinforcement
However, over time and under certain conditions like carbonation, chloride ingress, and cracking this protective state can be compromised. Steel has areas with...
Reinforcement Schedules
Once a behavior is learned,...

