Related Experiment Video
Updated: Sep 10, 2025

Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit
Published on: September 8, 2023
Decentralized queue control with delay shifting in edge-IoT using reinforcement learning
1Vinnytsia National Technical University, Vinnytsia, Ukraine. kovtun_v_v@vntu.edu.ua.
Abstract:
The article presents an adaptive approach to modelling and managing the service process of requests at peripheral nodes of edge-IoT systems. This approach is highly relevant in light of increasing demands for energy efficiency, responsiveness, and self-regulation under unstable traffic conditions. A stochastic G/G/1 model with a parameterised time shift is proposed, accounting for the temporary unavailability of the device prior to request processing. Analytical expressions for key QoS indicators (delay, variability, loss, energy consumption) as functions of the shift parameter are derived, and a multi-factor reward function is constructed. A DQN-based reinforcement learning agent architecture is implemented to dynamically control the shift parameter in a decentralised manner based on the local real-time queue state. Experimental results using real-world datasets demonstrated a reduction in average delay by 17-26%, decreased fluctuations in service time, and improved queue recovery stability after peak loads compared to current state-of-the-art models. The proposed solution is traffic-type agnostic and scalable across edge architectures of varying complexity. The results are suitable for deployment in sensor networks, 5G/6G edge scenarios, and systems with dynamic QoS and energy management.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Distributed Loads: Problem Solving
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Feedback control systems
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...

