通过强化学习的马尔科夫跳跃系统的多人差分游戏
IEEE transactions on cybernetics
|March 3, 2025
概括
本研究介绍了马尔科夫跳跃系统 (MJS) 中多人差异游戏 (MDGs) 的新增强化学习 (RL) 方法. 该研究开发了分布式minmax和异步纳什策略,用于在复杂系统中增强控制政策的优化.
科学领域:
- 控制理论 控制理论
- 游戏理论 游戏理论
- 机器学习 机器学习
背景情况:
- 马尔科夫跳跃系统 (MJSs) 的在线多人差异游戏 (MDGs) 提出了重要的控制挑战.
- 现有的强化学习 (RL) 算法通常需要完整的系统知识或同步更新.
研究的目的:
- 开发基于RL的新策略,以解决MJSs的千年发展目标.
- 解决分布式和异步设置中现有的RL算法的局限性.
主要方法:
- 建议使用分布式游戏代数里卡蒂方程 (DGAREs) 和一种新的在线分布式RL算法进行分布式minmax策略.
- 开发一个在线异步RL算法,用于纳什战略在千年发展目标中的应用.
- 严格分析拟议的RL算法的融合.
主要成果:
- 分布式的minmax策略允许玩家在没有对其他玩家的策略事先了解的情况下推导出最佳策略.
- 异步RL算法结合了用于政策评估和改进的最新信息.
- 拟议的方法通过对反向摆形系统的模拟来证明其有效性.
结论:
- 新型RL算法在分布式minmax和纳什策略下有效解决MJS的千年发展目标.
- 开发的方法为动态系统的去中心化控制和自适应性政策优化提供了进步.
相关概念视频
Multi-input and Multi-variable systems
93
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
93
Reinforcement
172
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
172
Reinforcement Schedules
126
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
126
Observational Learning
118
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
118
Collisions in Multiple Dimensions: Problem Solving
3.5K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
3.5K
Dynamic Equilibrium
49.9K
A reversible chemical reaction represents a chemical process that proceeds in both forward (left to right) and reverse (right to left) directions. When the rates of the forward and reverse reactions are equal, the concentrations of the reactant and product species remain constant over time and the system is at equilibrium. A special double arrow is used to emphasize the reversible nature of the reaction. The relative concentrations of reactants and products in equilibrium systems vary greatly;...
49.9K


