弹性多代理RL:为数据丢失的分布式环境引入DQ-RTS.
Lorenzo Canese1, Gian Carlo Cardarilli2, Luca Di Nunzio2
1Department of Electronics, University of Rome Tor Vergata, 00133, Rome, Italy. canese@ing.uniroma2.it.
Scientific reports
|January 23, 2024
概括
本文介绍了DQ-RTS,一个分散的多代理强化学习算法. 它在动态环境中提高了沟通和适应能力,在变化的代理人数量下显示了更快的融合和强大的性能.
科学领域:
- 计算机科学 计算机科学
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
背景情况:
- 分布式系统面临的挑战是不可靠的通信和动态的代理群体.
- 现有的多代理强化学习 (MARL) 算法在非理想的,波动的环境中扎.
研究的目的:
- 提出DQ-RTS,一个新的分散的MARL算法.
- 增强在分布式环境中的代理沟通和适应能力.
- 在具有挑战性的场景中,对现有方法进行DQ-RTS评估.
主要方法:
- 开发了具有优化通信协议的DQ-RTS.
- 使用Q-RTS (实时群体的Q学习) 进行了比较分析.
- 进行了广泛的实验,对具有不同代理人数量和通信质量的基准任务进行了广泛的实验.
主要成果:
- 在非理想的通信条件下,DQ-RTS表现出优越的融合速度,加速度系数为1.6至2.7.
- 算法保持了性能稳定性,尽管代理种群的波动.
- 在各种基准任务中验证了可扩展性和有效性.
结论:
- 在动态分布式环境中,DQ-RTS为MARL提供了一种实用且有弹性的解决方案.
- 该算法有效地解决了非理想的通信和不同的代理人数量.
- 与Q-RTS相比,DQ-RTS在融合和稳定性方面取得了显著的改进.
相关概念视频
Distribution Reliability and Automation
107
Distribution reliability in electrical power systems is critical for ensuring an uninterrupted power supply to consumers at minimal cost. According to IEEE Standard Terms, reliability is the probability that a device will function without failure over a specified time period or amount of usage. For electric power distribution, this translates to maintaining continuous power supply and addressing customer concerns over power outages. Several indices, as defined by IEEE Standard 1366-2012, are...
107
Reinforcement Schedules
148
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
148
Multi-input and Multi-variable systems
106
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
106
Reducing Line Loss
154
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
154
Randomized Experiments
6.9K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.9K
Associative Learning
370
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
370


