Related Experiment Video
Updated: Oct 10, 2025

05:30
Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit
Published on: September 8, 2023
728
Multi-Objective Optimization of Energy Saving and Throughput in Heterogeneous Networks Using Deep Reinforcement
1Department of Computer Engineering, Gachon University, Seongnam 13120, Korea.
Sensors (Basel, Switzerland)
|December 10, 2021
Summary
This study introduces a novel deep reinforcement learning algorithm for energy-efficient routing in dense wireless networks. The method optimizes energy savings and user throughput, achieving results comparable to traditional optimization methods.
Area of Science:
- Wireless communication networks
- Green networking and energy efficiency
- Optimization techniques
Background:
- Mobile service providers deploy small cells for enhanced wireless networking using mmWave frequencies.
- Denser small-cell heterogeneous networks (HetNets) face significant power consumption challenges.
- Green networking is crucial for reducing CO2 emissions and combating climate change.
Purpose of the Study:
- To investigate the feasibility of deep reinforcement learning (DRL) for energy-efficient routing and throughput maximization in HetNets.
- To develop a dual-objective optimization model addressing both energy consumption and user throughput.
- To explore a novel application of DRL in a previously unaddressed wireless networking problem.
Main Methods:
- Formulated a mixed integer linear programming (MILP) dual-objective optimization model.
- Proposed a proximal policy (PPO)-based multi-objective deep reinforcement learning algorithm.
- Utilized an actor-critic model within an optimistic linear support framework for iterative solution searching.
Main Results:
- The proposed DRL algorithm demonstrated comparable performance to CPLEX in achieving energy savings and maximizing user throughput.
- Experimental results validate the effectiveness of the PPO-based approach for dual-objective optimization in wireless networks.
- The algorithm successfully balances energy efficiency with user throughput in dense HetNets.
Conclusions:
- Deep reinforcement learning, specifically the PPO algorithm, is a viable and effective approach for energy-efficient routing and throughput maximization in wireless HetNets.
- The developed algorithm offers a near-optimal online solution with fast inference times.
- This research contributes to the advancement of green networking solutions for future wireless infrastructures.
Related Concept Videos
Distributed Loads: Problem Solving
813
Beams are structural elements commonly employed in engineering applications requiring different load-carrying capacities. The first step in analyzing a beam under a distributed load is to simplify the problem by dividing the load into smaller regions, which allows one to consider each region separately and calculate the magnitude of the equivalent resultant load acting on each portion of the beam. The magnitude of the equivalent resultant load for each region can be determined by calculating...
813
Reinforcement
423
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
423
Reinforcement Schedules
261
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
261
Multi-input and Multi-variable systems
195
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
195
Maximum Power Flow and Line Loadability
214
The maximum power flow for lossy transmission lines is derived using ABCD parameters in phasor form. These parameters create a matrix relationship between the sending-end and receiving-end voltages and currents, allowing the determination of the receiving-end current. This relationship facilitates calculating the complex power delivered to the receiving end, from which real and reactive power components are derived.
214
Observational Learning
372
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
372

