一个可扩展的方法,以优化交通信号控制与联合强化学习加强学习
Jingjing Bao1, Celimuge Wu2, Yangfei Lin1
1Department of Computer and Network Engineering, The University of Electro-Communications, Tokyo, 182-8585, Japan.
Scientific reports
|November 6, 2023
概括
本研究介绍了基于联合学习的强化学习,用于交通信号控制,改善交通流量和减少延误. 这种新的方法通过有效地整合本地十字路口数据来增强全球交通管理.
科学领域:
- 智能运输系统 智能运输系统
- 在交通管理中的人工智能.
背景情况:
- 交通信号控制 (TSC) 对于减轻城市拥堵,行程时间和环境影响至关重要.
- 对于TSC来说,传统的强化学习 (RL) 在集中和分布式系统中面临着沟通和适应的挑战.
- 深度学习和物联网的进步需要创新的TSC解决方案.
研究的目的:
- 为交通信号控制提出一种基于联邦学习 (FL) 的新型RL方法.
- 解决适应性TSC中集中和分布式RL的局限性.
- 创建一个统一的代理国家结构,以便在各个交叉点之间有效地整合知识.
主要方法:
- 实施联合学习 (FL),将当地交通信号代理人的知识整合到全球模型中.
- 使用统一的代理状态结构来克服交叉点之间的变化.
- 将RL神经网络的部分集成到云端,并在融合时微调剩余层.
主要成果:
- 显著减少了全球排队和等待时间.
- 在摩纳哥的真实交通网络上验证了模型的可扩展性.
- 展示了该模型的适应性和与新的交通交叉点集成的潜力.
结论:
- 基于联邦学习的RL提供了一个强大的解决方案,用于自适应的交通信号控制.
- 拟议的方法有效地平衡了全球交通优化与当地交叉路口特征.
- 这种方法有望提高智能运输系统的效率和可扩展性.
相关概念视频
Load-frequency control
169
Load-frequency control (LFC) is vital for maintaining power system stability, ensuring that frequency and power flows remain within acceptable limits during load changes. Turbine-governor control eliminates rotor accelerations and decelerations following load changes. However, a steady-state frequency error persists when the change in the turbine-governor reference setting is zero. In an interconnected power system, each area agrees to export or import a scheduled amount of power through...
169
Reinforcement Schedules
160
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
160
Feedback control systems
319
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
319
SFG Algebra
121
In Signal Flow Graph (SFG) algebra, the value a node represents is determined by the sum of all signals entering that node. This summed value is then transmitted through every branch leaving the node, making the SFG a powerful tool for visualizing and analyzing control systems.
Each node in an SFG corresponds to a variable, and the interactions between nodes are represented by branches with associated gains. When multiple branches lead into a node, the value at that node is the sum of the...
Each node in an SFG corresponds to a variable, and the interactions between nodes are represented by branches with associated gains. When multiple branches lead into a node, the value at that node is the sum of the...
121
Controller Configurations
102
Controller configurations are crucial in a car's cruise control system because they manage speed over time to maintain a consistent pace regardless of road conditions, thereby meeting design goals. In traditional control systems, fixed-configuration design involves predetermined controller placement. System performance modifications are known as compensation.
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
102
Root-Locus Method
160
A cruise control system in a car is designed to maintain a specified speed automatically by adjusting the gas pedal. The system continuously measures the vehicle's speed and makes fine adjustments to the pedal to achieve this goal. The root locus method is particularly useful for understanding how the cruise control system's behavior changes under varying conditions, such as when the car goes uphill, downhill, or faces strong wind resistance.
This system can be represented by a block...
This system can be represented by a block...
160


