基于多代理深度强化学习的空对空通信系统的多目标优化.
Shaofu Lin1, Yingying Chen1, Shuopeng Li1
1Faculty of Information Technology, Beijing University of Technology, Beijing 100124, China.
Sensors (Basel, Switzerland)
|December 9, 2023
概括
这项研究介绍了无人机的新型空对空通信系统,集成移动边缘计算和无线电力传输. 该系统优化了能源,并减少了扩展无人机任务的计算延迟.
科学领域:
- *专注于空中通信系统和移动边缘计算.
- * 集成无线电源传输,以提高无人机耐力.
背景情况:
- *无人驾驶飞行器 (UAV) 系统由于能源限制而面临飞行时间和计算能力的限制.
- * 现有的系统难以平衡实时数据处理和持续的空中操作.
研究的目的:
- *为无人机开发一个全双重的空对空通信系统 (A2ACS).
- *为了减少无人机的计算延迟和能源消耗,同时确保任务的连续性.
- *为了优化系统吞吐量,最大限度地减少低功耗警报,并高效地管理能量传输.
主要方法:
- * 提出了一种结合移动边缘计算和无线电力传输的新系统.
- * 开发了一个多目标优化框架,以平衡系统吞吐量,无人机功率状态,能源接收和能源消耗.
- *使用多代理深度决定性政策梯度 (MADDPG) 算法来优化AEES位置和能量传输功率.
- *使用K-means集群,以便在空边能源服务器 (AEES) 和无人机之间进行公平的关联.
主要成果:
- * 拟议的多目标DDPG (MODDPG) 算法与基线方法相比显示出更高的性能.
- * 该系统有效地减少了无人机的计算延迟和能源消耗.
- * 实现了系统吞吐量,无人机能量水平和能源服务器效率之间的平衡.
结论:
- * 集成的A2ACS模型提高了无人机的操作效率和耐久性.
- *基于MADDPG的优化为动态的空中环境提供了有效的决策.
- *K-means算法确保了公平的资源分配,提高了整体系统的公平性.
相关概念视频
Multi-input and Multi-variable systems
106
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
106
Reinforcement
212
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
212
Observational Learning
181
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
181
Reinforcement Schedules
149
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
149
Collisions in Multiple Dimensions: Problem Solving
4.2K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
4.2K
Masking and Demasking Agents
2.4K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.4K


