基于增强的深度强化学习的适应性城市交通信号控制
1School of Electronic Engineering, Xi'an Shiyou University, Xi'an, 710065, Shaanxi, China.
Scientific reports
|June 19, 2024
概括
这项研究通过使用深度强化学习 (DRL) 增强了交通信号控制 (TSC),提高了采样和稳定性. 新方法实现了更快的融合和更好的交通流管理在城市网络.
科学领域:
- 智能运输系统 智能运输系统
- 人工智能的人工智能
- 交通工程是交通工程.
背景情况:
- 目前的智能交通系统专注于交通信号控制 (TSC) 以提高城市效率.
- 深度增强学习 (DRL) 被广泛使用,但由于样本选择效率低下,它面临着缓慢融合的挑战,并需要在各种交通条件下提高稳定性.
研究的目的:
- 提出基于深度增强学习 (DRL) 的交通信号控制 (TSC) 的增强方法.
- 提高培训效率,模型稳定性和城市交通管理的整体性能.
主要方法:
- 利用决斗网络和双重Q学习来缓解DRL的高估问题.
- 实施了优先抽样机制,以有效地利用内存.
- 将噪声参数集成到神经网络中,以提高稳定性.
- 以矩阵形式表示交通数据,并采用循环阶段的行动空间.
- 设计了一个反映真实世界的交通场景的奖励功能.
主要成果:
- 与标准DRL方法相比,已经证明了更快的趋同.
- 在减少队列长度和等待时间方面实现了最佳性能.
- 实验结果证实了该方法在各种交通流动场景中的稳定性.
结论:
- 拟议的基于DRL的TSC方法在培训效率和绩效方面提供了显著的改进.
- 这些改进带来了更强大,更有效的信号控制,特别是在复杂的城市环境中.
相关概念视频
PD Controller: Design
218
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
218
Reinforcement
202
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
202
Real-World Application of Classical Conditioning
545
Classical conditioning not only includes the initial pairing of stimuli but also extends to more complex forms, such as higher-order conditioning. Higher-order conditioning involves creating associations beyond the primary conditioned stimulus, resulting in a chain of conditioned responses.
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
Higher-order, or second-order, conditioning occurs when a neutral stimulus becomes associated with an already established conditioned stimulus through repeated pairings. For instance, if a dog has been...
545


