Related Experiment Video
Updated: Jan 16, 2026

Automated Deployment of an Internet Protocol Telephony Service on Unmanned Aerial Vehicles Using Network Functions Virtualization
Published on: November 26, 2019
Decentralized resource allocation in UAV communication networks through reward based multi agent learning.
Muhammad Shoaib1, Ghassan Husnain2, Muhsin Khan3
1Department of Computer Science, CECOS University of IT and Emerging Sciences, Peshawar, 25100, Pakistan.
Unmanned aerial vehicles (UAVs) optimize wireless access by independently managing resources like users and power. A novel reward-based multi-agent learning (RMAL) framework maximizes system rewards without full information exchange.
Area of Science:
- Wireless Communication Systems
- Artificial Intelligence in Networking
- Resource Management
Background:
- Unmanned aerial vehicles (UAVs) offer flexible aerial base stations (ABS) for on-demand wireless access.
- Dynamic resource allocation is crucial for maximizing performance in multi-UAV systems.
- Existing methods may require extensive information exchange, increasing system overhead.
Purpose of the Study:
- To investigate dynamic resource allocation strategies for multi-UAV-enabled communication systems.
- To maximize the long-term rewards and overall system performance.
- To develop a decentralized learning framework for UAV resource management.
Main Methods:
- Modeling the resource allocation problem as a stochastic game with UAVs as learning agents.
- Developing a reward-based multi-agent learning (RMAL) framework.
- Implementing an agent-independent strategy using a Q-learning-based framework with local observations.
Main Results:
- The proposed RMAL framework achieves effective resource allocation without requiring full information exchange between UAVs.
- Simulation results indicate that RMAL performance is sensitive to parameter tuning but offers a good trade-off.
- The method provides acceptable performance compared to systems with complete information sharing.
Conclusions:
- The RMAL framework offers an efficient approach to dynamic resource allocation in multi-UAV systems.
- Decentralized learning with local observations can achieve near-optimal performance while reducing communication overhead.
- This research contributes to the advancement of intelligent resource management in aerial communication networks.
Related Concept Videos
Observational Learning
Distributed Loads: Problem Solving
Optimal Foraging
Multi-input and Multi-variable systems
In the absence of...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Absolute Motion Analysis- General Plane Motion
As the drone's propellers rotate, an upward force is generated that counteracts the force of gravity, enabling the drone to lift off from the ground. This initial movement of the drone is along a straight path, representing a form of translational motion. In this phase, every point on the...
