Related Experiment Video
Updated: Jul 18, 2025

Closed-loop Neuro-robotic Experiments to Test Computational Properties of Neuronal Networks
Published on: March 2, 2015
User Pairing for Delay-Limited NOMA-Based Satellite Networks with Deep Reinforcement Learning
Qianfeng Zhang1,2, Kang An3, Xiaojuan Yan1,4
1Guangxi Key Laboratory of Ocean Engineering Equipment and Technology, Qinzhou 535011, China.
This study introduces a deep reinforcement learning (DRL) approach for optimal user pairing in non-orthogonal multiple access (NOMA) satellite networks. The DRL method enhances resource utilization and performance for diverse Quality of Service (QoS) requirements.
Area of Science:
- Satellite Communications
- Wireless Networks
- Network Optimization
Background:
- Non-Orthogonal Multiple Access (NOMA) enables efficient resource sharing in satellite networks.
- Varying Quality of Service (QoS) requirements, particularly delay constraints, pose challenges for NOMA performance.
- Effective capacity is a key metric for evaluating performance under delay-limited QoS.
Purpose of the Study:
- To address the user pairing problem in power-domain NOMA satellite networks with diverse delay QoS requirements.
- To optimize resource utilization and network performance by efficiently forming NOMA user pairs.
- To develop a dynamic user selection strategy that overcomes the intractability of traditional methods.
Main Methods:
- Utilized effective capacity to model the impact of delay QoS on performance.
- Determined a power allocation coefficient to maintain performance for delay-sensitive users compared to Orthogonal Multiple Access (OMA).
- Employed a Deep Reinforcement Learning (DRL) algorithm for dynamic and optimal user selection in delay-limited NOMA scenarios.
Main Results:
- The DRL algorithm successfully identified optimal user pairings for NOMA in each state.
- The proposed DRL-based scheme demonstrated superior performance compared to random selection and OMA.
- The approach effectively managed channel conditions and delay QoS requirements for user pairing.
Conclusions:
- The DRL-based user selection scheme provides an effective solution for optimizing NOMA satellite networks.
- This method significantly improves performance by dynamically pairing users based on QoS and channel conditions.
- The findings highlight the potential of DRL in advancing complex wireless network management.
Related Concept Videos
Reinforcement Schedules
Once a behavior is learned,...
Observational Learning
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Associative Learning
Classical conditioning, also known...
Real-World Application of Classical Conditioning
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...

