Related Experiment Video
Updated: Jan 9, 2026

An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
Published on: August 2, 2018
Higher-Order Markov Model-Based Analysis of Reinforcement Learning in 6G Mobile Retrial Queueing Systems
1Faculty of Informatics, University of Debrecen, Egyetem ter 1, 4032 Debrecen, Hungary.
Deep Q-Network Reinforcement Learning (DQN-RL) optimizes 6G mobile networks. Markov chain analysis shows 5-10 training episodes are sufficient for efficient policy convergence, enhancing performance and reducing energy use.
Area of Science:
- Telecommunications Engineering
- Artificial Intelligence
- Queueing Theory
Background:
- 6G mobile communication services face dynamic challenges in queueing systems.
- Deep Q-Network Reinforcement Learning (DQN-RL) offers a potential solution for optimizing network behavior.
- Understanding agent learning convergence is crucial for effective implementation.
Purpose of the Study:
- To analyze the learning convergence of DQN-RL agents in 6G retrial queueing systems.
- To quantify convergence characteristics using Markov chain methods and mixing time analysis.
- To provide a foundation for optimizing 6G queueing strategies under uncertainty.
Main Methods:
- Utilized first- and second-order Markov chain methods to analyze DQN-RL agent convergence.
- Simulated temporal evolution of reward sequences as Markov chains.
- Employed mixing time analysis and spectral gap properties of Markov models to assess convergence.
Main Results:
- Markov chain analysis indicates 10 training episodes are sufficient for policy convergence in DQN-RL.
- In some scenarios, as few as 5 episodes enhance mobile network performance with low energy consumption.
- Mixing time calculations assessed learning stability and system responsiveness across 120 parameter combinations.
Conclusions:
- DQN-RL convergence is sensitive to system parameters and retrial dynamics in 6G queueing.
- Markov chain analysis provides a rigorous method for evaluating learning convergence.
- The findings support optimizing 6G queueing strategies for enhanced efficiency and performance.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Parameters Affecting Nonlinear Elimination: Zero-Order Input, First-Order Absorption and Two-Compartment Model
When a drug is administered through a constant intravenous infusion and eliminated via nonlinear pharmacokinetics, it follows zero-order input. For example, oral drugs undergo first-order absorption upon administration and are eliminated through nonlinear pharmacokinetics.
In the case of subcutaneously administered drugs,...
Current Growth And Decay In RL Circuits
Multi-input and Multi-variable systems
In the absence of...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Multimachine Stability
In analyzing the system, the nodal equations represent the relationship between bus voltages, machine voltages, and machine currents. The nodal equation is given by:

