Related Experiment Video
Updated: Jun 29, 2026

Evaluation of an Exclusive Spur Dike U-Turn Design with Radar-Collected Data and Simulation
Published on: February 1, 2020
Optimization of urban freight intelligent route based on reinforcement learning
Guizhe Xin1,2, Yuqing Tang3, Hongzhen Gao4
1College of Architecture and Urban Planning, Tongji University, Shanghai, 200092, China.
This study introduces a new AI framework, constraint-aware projected policy learning-reinforcement learning (CAPPL-RL), to optimize urban freight delivery routes. CAPPL-RL significantly reduces delivery times, fuel consumption, and constraint violations in dynamic city environments.
Area of Science:
- Artificial Intelligence
- Operations Research
- Logistics and Supply Chain Management
Background:
- Urban freight systems face inefficiencies due to traffic congestion, delivery time constraints, and diverse vehicle fleets.
- Traditional routing methods fail to adapt to real-time, partially observable, multi-constrained conditions, leading to increased delays and costs.
Purpose of the Study:
- To develop and validate a novel reinforcement learning framework, CAPPL-RL, for optimizing urban freight delivery routes.
- To incorporate real-world constraints such as traffic, time windows, and vehicle capacity into the routing optimization process.
Main Methods:
- Formulated the urban freight routing problem as a partially observable Markov decision process.
- Developed a constraint-aware projected policy learning-reinforcement learning (CAPPL-RL) framework integrating Q-learning, ε-greedy exploration, and projection-based constrained policy optimization (PCPO).
- Utilized the UFVOD dataset for simulation and validation.
Main Results:
- CAPPL-RL reduced average delivery time by 20.2% (from 65.3 to 52.1 min).
- Fuel consumption decreased by 22.5% (from 0.12 to 0.093 L/km).
- Time window constraint violations were reduced by 75% (from 12 to 3), with 100% vehicle capacity compliance.
Conclusions:
- CAPPL-RL demonstrates superior performance compared to existing methods like PCPO-RL in dynamic urban freight logistics.
- The framework is robust, scalable, and adaptive, offering significant improvements in efficiency and constraint adherence for urban delivery operations.
Related Concept Videos
Optimal Foraging
Short-distance Transport of Resources
Distributed Loads: Problem Solving
Rolling Resistance: Problem Solving
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Reinforcement Schedules
Once a behavior is learned,...
