在无人机支持的物联网网络中,用于空地通信规划的多代理DRL
Khalid Ibrahim Qureshi1, Bingxian Lu1, Cheng Lu1
1Key Laboratory for Ubiquitous Network and Service Software of Liaoning Province, School of Software, Dalian University of Technology, Dalian 116024, China.
Sensors (Basel, Switzerland)
|October 26, 2024
概括
这项研究通过分离上联和下联用户协会来增强无人机通信网络. 这种新的方法优化了无人机轨迹和用户连接,以提高网络效率,特别是在紧急情况下.
科学领域:
- 无线通信无线通信
- 网络工程 网络工程
- 机器人技术 机器人技术 机器人技术
背景情况:
- 现有的无人驾驶飞行器 (UAV) 辅助网络将上链和下链联结起来,导致在动态环境中性能不佳.
- 无法预测的用户需求和网络条件挑战了传统无人机通信系统的效率.
研究的目的:
- 为了提高UAV辅助通信网络的总和率效率.
- 为了提高网络效率,将地面用户 (GBU) 的上联和下联协会分开.
- 整合无人机轨迹设计和用户协会,以最大限度地提高网络总率效率.
主要方法:
- 为无人机轨迹和用户协会制定了一个全面的优化问题.
- 将非凸问题改制为部分可观测的马尔科夫决策过程 (POMDP).
- 采用多代理深度强化学习 (MADRL),特别是多代理深度决定性政策梯度 (MADDPG) 算法.
主要成果:
- 通过解的上链和下链协会实现了增强的总额率效率.
- 使无人机能够通过POMDP使用本地观测进行实时决策.
- 通过使用MADDPG证明了最佳用户关联和轨迹控制的有效学习.
结论:
- 拟议的混合MADRL框架平衡了集中训练与分布式执行,以实现最佳的无人机操作.
- 解关联方法在动态环境中显著提高了网络效率.
- 该解决方案非常适用于危机应对和搜救任务等关键场景.
相关概念视频
Reinforcement
181
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
181
Masking and Demasking Agents
2.4K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.4K
Observational Learning
142
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
142
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
Air-entraining Agents
76
Air-entraining agents improve the durability and workability of concrete in climates with frequent freezing and thawing. These agents prevent cracks by introducing small air bubbles into the mix, creating spaces accommodating water expansion when temperatures drop. The air-entraining agents lower the surface tension of water, forming stable, small air bubbles. This method is more effective than having accidental large voids, as the intentional, smaller, and evenly distributed air voids improve...
76
Cognitive Learning
222
Cognitive learning is based on purposive behavior, incidental learning, and insight learning.
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
E. C. Tolman's theory of purposive behavior emphasizes that much behavior is goal-directed. He argued that to understand behavior, we must look at the entire sequence of actions leading to a goal. For instance, high school students study hard, not just due to past reinforcement but also to achieve the goal of getting into a good college.
Tolman introduced the idea that behavior is influenced by...
222


