在线强化学习控制设计与加速机制的未知多代理系统通过价值代的加速机制
概括
本研究介绍了一种在线强化学习 (RL) 控制方法,用于未知的多元代理系统 (MAS). 这种方法确保了稳定性,并通过事件触发机制和神经网络加速了趋同.
科学领域:
- 控制系统工程 控制系统工程
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 多代理系统 (MAS) 存在复杂的控制挑战,特别是当系统动态未知时.
- 最佳的合作控制对于MAS的协调行为至关重要.
- 在不断变化的控制政策下,现有的方法往往难以保证稳定性.
研究的目的:
- 为未知的线性离散时间多代理系统 (MAS) 开发在线强化学习 (RL) 控制方法.
- 确保MAS的稳定性,即使在学习过程中产生不成熟的政策.
- 为了加快价值代 (VI) 算法的收率.
主要方法:
- 设计了一个具有不断变化的政策的在线学习计划,其中包含了一个事件触发机制来过可接受的控制政策.
- 开发了一个稳定性标准,以保证系统的稳定性,而不需要单调的值函数序列.
- 引入了VI的加速机制,分析了放松因子对收速度的影响.
- 逆向传播 (BP) 神经网络 (NN) 被用于实际实施.
主要成果:
- 提出的方法成功地解决了未知线性离散时间MAS的最佳合作控制问题.
- 事件触发稳定性标准有效过政策,确保系统稳定性.
- 加速机制显著提高了VI算法的收率.
- 模拟结果验证了开发的控制方法的有效性和性能.
结论:
- 开发的在线RL控制方法为在未知的MAS中进行合作控制提供了强大的解决方案.
- 事件触发机制和加速技术的整合提供了更好的稳定性和效率.
- 使用BP神经网络证明了拟议方法的实际可用性.
相关概念视频
Multi-input and Multi-variable systems
90
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
90
Feedback control systems
256
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
256
Reinforcement Schedules
119
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
119
Open and closed-loop control systems
577
Control systems are foundational elements in automation and engineering. They are broadly categorized into open-loop and closed-loop systems. These classifications hinge on the presence or absence of feedback mechanisms, significantly influencing the system's performance, complexity, and application.
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
577
Linear Momentum in Control Volume
653
Newton's second law is applied to obtain the linear momentum in a control volume in a fluid system. According to this law, the rate of change of linear momentum is equal to the sum of external forces acting on the system. When a control volume matches the fluid system at a specific moment, the forces acting on both are identical. Reynolds transport theorem helps explain this by breaking down the system's linear momentum into two components: the rate of change of linear momentum within...
653
Three-Dimensional Force System:Problem Solving
586
A three-dimensional force system refers to a scenario in which three forces act simultaneously in three different directions. This type of problem is commonly encountered in physics and engineering, where it is necessary to calculate the resultant force on the system, which can then be used to predict or analyze the behavior of the object or structure under consideration.
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
586


