信用分配与预测性贡献测量在多代理强化学习的多代理强化学习
1School of Intelligence Science and Technology, Peking University, Beijing, 100871, China.
概括
本研究介绍了预测性贡献测量,这是一种用于多代理强化学习的新型学分分配方法. PC-MAPPO增强了政策梯度方法,在复杂的合作任务中表现优于现有的方法.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 强化学习是一种强化学习.
背景情况:
- 在使用集中培训和分散执行的多代理系统中,信用指派具有挑战性.
- 价值分解方法在Q学习中工作得很好,但与政策梯度方法相斗争.
- 现有的方法在复杂的多代理情景中有效地分配信贷方面存在局限性.
研究的目的:
- 引入预测性贡献测量 (PCM),用于多代理任务的明确信用分配方法.
- 通过将PCM与MAPPO集成,开发预测性贡献多代理近接政策优化 (PC-MAPPO).
- 为拟议的信用分配机制提供理论担保.
主要方法:
- 通过比较代理预测错误,开发了预测性贡献测量 (PCM).
- 根据代理人对全球状态过渡的相关性分配了代用奖励.
- 将PCM集成到多代理近接政策优化 (MAPPO) 框架中,创建PC-MAPPO.
- 利用预先训练的预测器来提高性能.
主要成果:
- 在StarCraft的多代理挑战地图上,PC-MAPPO在MAPPO,QMIX和加权QMIX上表现优越.
- 在需要高水平合作的地图上观察到显著的性能增长.
- 在并行培训中,PC-MAPPO取得了最先进的结果,提高了数据效率.
结论:
- 预测性贡献测量为基于政策梯度的多代理强化学习中的信用分配提供了有效的解决方案.
- PC-MAPPO显著提高了合作任务的性能,特别是在具有挑战性的场景中.
- 拟议的方法对推进多代理强化学习研究和应用有前途.
相关概念视频
Multi-input and Multi-variable systems
134
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
134
Observational Learning
246
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
246
Predicting Reaction Outcomes
8.5K
Kinetics describes the rate and path by which a reaction occurs. In contrast, thermodynamics deals with state functions and describes the properties, behavior, and components of a system. It is not concerned with the path taken by the process and cannot address the rate at which a reaction occurs. Although it does provide information about what can happen during a reaction process, it does not describe the detailed steps of what appears on an atomic or a molecular level. On the other hand,...
8.5K
Associative Learning
465
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
465
Attribution Theory
13.1K
Behavior is a product of both the situation (e.g., cultural influences, social roles, and the presence of bystanders) and of the person (e.g., personality characteristics). Subfields of psychology tend to focus on one influence or behavior over others. Situationism is the view that our behavior and actions are determined by our immediate environment and surroundings. In contrast, dispositionism holds that our behavior is determined by internal factors (Heider, 1958).
13.1K
Reinforcement
295
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
295


