Related Experiment Video
Updated: Aug 16, 2025

05:30
Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit
Published on: September 8, 2023
634
Joint Deep Reinforcement Learning and Unsupervised Learning for Channel Selection and Power Control in D2D Networks
Ming Sun1, Yanhui Jin1, Shumei Wang2
1College of Computer and Control Engineering, Qiqihar University, Qiqihar 161006, China.
Entropy (Basel, Switzerland)
|December 23, 2022
Summary
This study introduces a novel distributed resource allocation algorithm for device-to-device (D2D) communications. The algorithm uses deep Q-networks and unsupervised learning to reduce interference and maximize spectrum utilization in 5G networks.
Area of Science:
- Wireless Communication
- Network Resource Management
- Artificial Intelligence in Telecommunications
Background:
- Device-to-device (D2D) communication offers a solution to spectrum scarcity in 5G networks.
- Shared channels in D2D networks lead to significant interference, hindering capacity and spectral efficiency.
- Existing resource allocation methods often rely on centralized control, requiring extensive network information.
Purpose of the Study:
- To develop a distributed resource allocation algorithm for D2D networks.
- To mitigate interference and enhance wireless spectrum utilization.
- To improve network capacity and spectral efficiency in 5G systems.
Main Methods:
- A deep Q-network (DQN) was employed for distributed channel allocation in dynamic environments.
- An unsupervised learning-based deep neural network was utilized for optimized power control.
- The proposed algorithm operates distributively, using local information for channel selection and power control.
Main Results:
- The proposed algorithm effectively reduces interference in D2D user pairs.
- It significantly enhances network capacity and wireless spectrum utilization.
- Simulation results demonstrate superior performance compared to traditional centralized and distributed algorithms in terms of convergence speed and transmit sum-rate maximization.
Conclusions:
- The integration of DQN and unsupervised learning provides an effective distributed solution for D2D resource allocation.
- The algorithm achieves higher spectral efficiency and faster convergence than conventional methods.
- This approach offers a scalable and efficient strategy for future wireless communication systems.
Related Concept Videos
Maximum Power Transfer
330
Numerous practical applications within engineering disciplines, such as telecommunications, necessitate optimizing power delivery to a connected load. This pursuit, however, entails inherent internal losses, which can either equal or exceed the power supplied to the load. The Thevenin equivalent circuit is helpful in finding the maximum power a linear circuit can deliver to a load. It is assumed in this context that the load resistance can be adjusted.
By substituting the entire circuit with...
By substituting the entire circuit with...
330
Maximum Power Flow and Line Loadability
158
The maximum power flow for lossy transmission lines is derived using ABCD parameters in phasor form. These parameters create a matrix relationship between the sending-end and receiving-end voltages and currents, allowing the determination of the receiving-end current. This relationship facilitates calculating the complex power delivered to the receiving end, from which real and reactive power components are derived.
158
Observational Learning
260
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
260
Associative Learning
503
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
503
Reinforcement
308
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
308
Reinforcement Schedules
227
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
227

