基于强化学习的漏斗控制和隐私保护,用于具有输入死区输入的多代理系统
Jiaxin Huang1, Xiaoyang Liu1, Sikai Shen1
1School of Computer Science and Technology, Jiangsu Normal University, Xuzhou, 221116, Jiangsu, PR China.
概括
本研究介绍了一种保护隐私的适应式漏斗控制器,用于具有输入死区约束的多代理系统,确保准确的跟踪,同时使用强化学习和加密技术保护状态信息.
科学领域:
- 控制系统工程 控制系统工程
- 人工智能的人工智能
- 网络安全 网络安全
背景情况:
- 多代理系统 (MAS) 经常面临输入死区约束和通信负担的挑战.
- 在分布式控制系统中,在国家信息传输过程中确保数据隐私至关重要.
- 需要适应性控制策略来处理复杂系统中的不确定性和非线性.
研究的目的:
- 为MAS设计一个保护隐私的基于强化学习的漏斗控制器,具有输入死区约束.
- 为了保证跟踪错误保持在规定的范围内,尽管系统的不确定性.
- 开发一个高效,安全的数据交换机制,用于MAS控制.
主要方法:
- 使用演员-关键强化学习模式制定了一个自适应的漏斗控制器.
- 模糊逻辑被用来近似未表征的系统非线性.
- 引入了一个事件触发的方案,以有效地更新控制信号.
- 为了安全的数据交换,Paillier加密方案被整合进来.
主要成果:
- 拟议的控制器成功地保证了跟踪错误在预定义的范围内.
- 事件触发方案有效地减少了通信负载,同时保持了控制性能.
- 密码机制确保了国家信息在传输过程中的隐私.
- 模拟验证了控制器在处理输入死区约束方面的可行性和有效性.
结论:
- 开发的战略为控制具有输入死区约束的多代理系统提供了强大而安全的解决方案.
- 强化学习,模糊逻辑和密码学的整合提高了系统性能和数据安全性.
- 这种方法为复杂的分布式系统中的高级,隐私意识的控制提供了基础.
相关概念视频
Multi-input and Multi-variable systems
385
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
385
Masking and Demasking Agents
3.4K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
3.4K
Reinforcement
826
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
826
Feedback control systems
685
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
685
Open and closed-loop control systems
1.6K
Control systems are foundational elements in automation and engineering. They are broadly categorized into open-loop and closed-loop systems. These classifications hinge on the presence or absence of feedback mechanisms, significantly influencing the system's performance, complexity, and application.
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
An open-loop control system operates without feedback from the output. It consists of two primary elements: the controller and the controlled process. The controller receives an input signal...
1.6K
Avoidance Learning and Learned Helplessness
2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K


