通过内部强化Q学习对受约束的多代理系统进行事件触发的最佳双边共识控制.
IEEE transactions on cybernetics
|June 24, 2025
概括
本研究介绍了针对多代理系统 (MAS) 的事件触发的最佳双方共识控制. 这种新的算法节省了资源,同时保证了稳定性,并提高了未知模型和输入和度的系统的性能.
科学领域:
- 控制理论 控制理论
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
背景情况:
- 调查二级离散时间多代理系统 (MAS) 的共识控制.
- 解决控制输入和和未知系统模型的挑战.
- 突出了传统的时间触发控制方法的局限性.
研究的目的:
- 为MAS开发一个事件触发的最佳双边共识控制策略.
- 提高MAS的学习效率和资源节约.
- 在控制约束下,以确保非对称稳定性和边界追踪误差.
主要方法:
- 定义了一个具有非二次函数的即时奖励信号来处理控制输入和.
- 引入了一种新的内部强化奖励功能,用于内在学习.
- 开发了一个数据驱动,事件触发的内部强化Q学习 (IrQL) 算法,使用演员关键神经网络.
主要成果:
- 拟议的事件触发的IrQL算法有效地利用环境,同时节省计算和传输资源.
- 功能分析和利亚普诺夫稳定理论证实了局限的内部强化奖励函数和MAS的非对称稳定性.
- 使用强化-关键行为体神经网络的在线实施表明了与现有方法相比的趋同和优越性能.
结论:
- 事件触发的IrQL算法提供了一个有效的解决方案,用于在MAS中以和和未知模型为最佳的双方共识控制.
- 这种方法显著提高了资源效率和系统稳定性.
- 模拟结果验证了算法的有效性和性能优势.
相关概念视频
Reinforcement
353
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
353
Reinforcement Schedules
243
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
243
Stability of Equilibrium Configuration: Problem Solving
678
The stability of equilibrium configurations is an important concept in physics, engineering, and other related fields. In simple terms, it refers to the tendency of an object or system to return to its equilibrium position after being disturbed. The stability of an equilibrium configuration can be analyzed by considering the potential energy function of the system and examining its behavior near the equilibrium point.
Problem-solving in the context of the stability of equilibrium configuration...
Problem-solving in the context of the stability of equilibrium configuration...
678
Multi-input and Multi-variable systems
152
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
152
Observational Learning
321
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
321
Open and closed-loop control systems
1.0K
Control systems are foundational elements in automation and engineering. They are broadly categorized into open-loop and closed-loop systems. These classifications hinge on the presence or absence of feedback mechanisms, significantly influencing the system's performance, complexity, and application.
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
1.0K


