Related Experiment Video
Updated: Sep 30, 2025

Integration of 5G Experimentation Infrastructures into a Multi-Site NFV Ecosystem
Published on: February 3, 2021
Multi-Agent Reinforcement Learning Based Fully Decentralized Dynamic Time Division Configuration for 5G and B5G
Xiangyu Chen1, Gang Chuai1, Weidong Gao1
1Department of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China.
Future network services require dynamic traffic adaptation. This study introduces a multi-agent reinforcement learning method for dynamic time division duplex configuration in 5G networks, optimizing uplink and downlink traffic efficiently.
Area of Science:
- Telecommunications Engineering
- Artificial Intelligence
- Network Optimization
Background:
- Future network services demand adaptive uplink and downlink traffic management.
- 5G New Radio (NR) requires dynamic adjustment of time-domain duplex patterns.
- Configuring effective dynamic time division duplex (D-TDD) patterns for 5G NR remains an open research challenge.
Purpose of the Study:
- To propose a decentralized D-TDD configuration method using distributed multi-agent deep reinforcement learning (MARL).
- To maximize the sum rates of all users (UE) by optimizing the D-TDD configuration policy.
- To reduce signaling overhead through a fully decentralized MARL approach.
Main Methods:
- Modeling the D-TDD configuration as a dynamic programming problem and defining the policy as a conditional probability distribution.
- Implementing a decentralized MARL solution where each base station (BS) acts as an agent, using local buffer length observations.
- Integrating a leniency controller and a binary LSTM (BLSTM) based auto-encoder to address MARL's global information limitations and handle complex data.
Main Results:
- The proposed distributed MARL method achieves stable convergence across diverse network environments.
- The system effectively configures uplink and downlink time slot ratios based on local user queue buffer lengths.
- The MARL approach, deployed on Mobile Edge Computing (MEC) servers, demonstrates superior performance compared to traditional distributed deep reinforcement algorithms.
Conclusions:
- The developed distributed MARL framework provides an effective and stable solution for dynamic TDD configuration in 5G networks.
- Decentralized decision-making based on local observations, enhanced by leniency control and auto-encoders, optimizes network resource allocation.
- This approach offers a promising direction for enhancing the adaptability and efficiency of future wireless communication systems.
More Related Videos
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Transformers in Distribution System
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
Distributed Loads: Problem Solving
Sequence Networks of Rotating Machines
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...

