基于强化学习的多代理系统的模糊双边共识:一种新的扩展政策之外的学习计划
IEEE transactions on cybernetics
|June 4, 2025
概括
本研究使用新的游戏理论方法解决了非线性多代理系统中的双边共识. 一个新的算法解决了复杂的方程,使得分布式控制,即使在未知的系统动态.
科学领域:
- 控制理论 控制理论
- 人工智能的人工智能
- 系统工程 系统工程
背景情况:
- 非线性多代理系统 (NMAS) 在实现协调行为方面存在挑战.
- 对于NMAS的分布式控制通常需要对系统动态的完整知识,这往往是不可用的.
- 双边共识 (BC) 是NMAS特定网络结构的关键目标.
研究的目的:
- 为了研究未知系统动态的非线性多代理系统 (NMAS) 的双边共识 (BC) 问题.
- 开发一种不依赖于对系统动态的明确知识的分布式控制策略.
- 将BC问题重新构成一个可解决的零和游戏.
主要方法:
- 使用Takagi-Sugeno (T-S) 模糊模型表示NMAS动态.
- 引入一个minmax游戏政策,以实现分布式控制.
- 重构BC问题作为一个零和游戏,可以通过游戏代数里卡蒂方程 (GAREs) 解决.
- 提出一种新的非政策代 (PI) 算法,以解决没有系统动态或初始政策的GARE.
主要成果:
- 拟议的扩展PI算法在学习过程中放松了对系统动态的依赖.
- 该算法消除了对初始可接受控制策略的需求,与传统的PI方法不同.
- 与标准值代技术相比,可以实现更快的融合速度.
- 该方法的有效性通过模拟和比较实验来证明.
结论:
- 开发的方法有效地解决了对NMAS具有未知的动态的双边共识问题.
- 新型缩放PI算法在数据需求和趋同速度方面提供了优势.
- 这项工作为复杂的多代理系统的分布式控制设计提供了强大的框架.
相关概念视频
Multi-input and Multi-variable systems
152
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
152
Observational Learning
321
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
321
Associative Learning
605
Associative learning is a fundamental concept in behavioral psychology, wherein a connection is established between two stimuli or events, leading to a learned response. This process is critical in understanding how behaviors are acquired and modified. Conditioning, the mechanism through which associations are formed, can be divided into two main types: classical conditioning and operant conditioning, each elucidating different aspects of associative learning.
Classical conditioning, also known...
Classical conditioning, also known...
605
Multicompartment Models: Overview
262
Multicompartment models are mathematical constructs that depict how drugs are distributed and eliminated within the body. They segment the body into several compartments, symbolizing various physiological or anatomical areas connected through drug transfer processes such as absorption, metabolism, distribution, and elimination.
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
262
Reinforcement
354
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
354
Avoidance Learning and Learned Helplessness
1.9K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.9K


