Related Experiment Video
Updated: May 16, 2025

14:55
Evaluating the Effect of Roadside Parking on a Dual-Direction Urban Street
Published on: January 20, 2023
3.2K
Federated deep reinforcement learning-based urban traffic signal optimal control
Mi Li1,2, Xiaolong Pan3, Chuhui Liu4
1College of Information Science and Engineering, Jiaxing University, Jiaxing, 314000, China. limi@zjxu.edu.cn.
Scientific Reports
|April 5, 2025
Summary
This study introduces a federated Proximal-Policy Optimization (PPO) method for intelligent traffic signal control, enhancing learning speed and model generalization in complex traffic networks while ensuring data privacy.
Area of Science:
- Artificial Intelligence
- Transportation Engineering
- Distributed Systems
Background:
- Deep reinforcement learning (RL) faces challenges in cross-domain traffic signal control, including slow learning and poor generalization.
- Existing methods struggle with non-independent and heterogeneous environmental data in real-world traffic scenarios.
Purpose of the Study:
- To develop a cross-domain intelligent traffic signal control method using federated Proximal-Policy Optimization (PPO).
- To address slow learning speeds and enhance model generalization for traffic signal optimization.
- To ensure information security and data privacy during distributed joint training.
Main Methods:
- Proposed a federated PPO approach for distributed joint training of agents across different traffic domains.
- Designed state, action, and reward functions, and optimized federated collaboration parameters.
- Implemented a novel state interaction method and reward function to improve agent collaboration and communication efficiency.
Main Results:
- Significantly reduced average vehicle waiting time by up to 27.34% compared to fixed timing.
- Achieved convergence speeds up to 47.69% faster than individual PPO and 45.35% faster than aggregated PPO.
- Demonstrated excellent robustness across various traffic flow settings.
Conclusions:
- The federated PPO method effectively optimizes intersection access efficiency and traffic flow.
- The approach enhances learning efficiency, fast convergence, and model generalization in cross-domain traffic control.
- Ensures information security and data privacy while improving communication efficiency in federated learning for traffic systems.
Related Concept Videos
PD Controller: Design
154
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
154
Feedback control systems
256
Feedback control systems are categorized in various ways based on their design, analysis, and signal types.
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
Linear feedback systems are theoretical models that simplify analysis and design. These systems operate under the principle that their output is directly proportional to their input within certain ranges. For instance, an amplifier in a control system behaves linearly as long as the input signal remains within a specific range. However, most physical systems exhibit inherent nonlinearity...
256
Reinforcement Schedules
119
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
119
Controller Configurations
75
Controller configurations are crucial in a car's cruise control system because they manage speed over time to maintain a consistent pace regardless of road conditions, thereby meeting design goals. In traditional control systems, fixed-configuration design involves predetermined controller placement. System performance modifications are known as compensation.
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
Control-system compensation involves various configurations, most commonly series or cascade compensation, in which the controller...
75
Multi-input and Multi-variable systems
91
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
In the absence...
91
Load-frequency control
97
Load-frequency control (LFC) is vital for maintaining power system stability, ensuring that frequency and power flows remain within acceptable limits during load changes. Turbine-governor control eliminates rotor accelerations and decelerations following load changes. However, a steady-state frequency error persists when the change in the turbine-governor reference setting is zero. In an interconnected power system, each area agrees to export or import a scheduled amount of power through...
97

