Related Experiment Video
Updated: Jun 4, 2025

14:23
Design and Optimization Strategies of a High-Performance Vented Box
Published on: June 9, 2023
1.1K
Carbon-Efficient Scheduling in Fresh Food Supply Chains with a Time-Window-Constrained Deep Reinforcement Learning
Yuansu Zou1,2, Qixian Gao1, Hao Wu1
1University of Electronic Science and Technology of China, Chengdu 611731, China.
Sensors (Basel, Switzerland)
|December 17, 2024
Summary
This study optimizes fresh food distribution routes using intelligent transportation systems (ITSs) and Internet of Things (IoT) to minimize costs and carbon emissions. A reinforcement learning model effectively manages logistics, considering time windows and cooling needs.
Area of Science:
- Intelligent Transportation Systems (ITSs)
- Supply Chain Management
- Operations Research
Background:
- Intelligent Transportation Systems (ITSs) integrate Internet of Things (IoT) for enhanced vehicle, infrastructure, and user connectivity, optimizing traffic flow.
- Fresh food supply chains face challenges in distribution, including time sensitivity, temperature control, and environmental impact.
- Minimizing logistics costs and carbon emissions is crucial for sustainable fresh food distribution.
Purpose of the Study:
- To develop an optimization model for fresh food supply chain distribution routes.
- To minimize total distribution costs, including carbon emission costs and cooling expenses.
- To enhance decision-making for optimizing transportation and distribution of fresh products.
Main Methods:
- Constructed an optimization model incorporating carbon taxes for emission costs, time windows, and cooling expenses.
- Utilized a graph attention network to represent node locations, paths, and data collection windows for path planning.
- Integrated a time-window-constrained reinforcement learning model to solve for optimal distribution routes.
Main Results:
- The proposed model effectively optimizes distribution routes for fresh food products.
- Demonstrated significant reductions in logistics costs and carbon emissions.
- Provided effective decision-making information for supply chain management under varying temperature conditions.
Conclusions:
- The time-window-constrained reinforcement learning model offers an effective solution for optimizing fresh food distribution.
- The model successfully balances cost reduction, time constraints, and environmental sustainability.
- ITSs and IoT technologies are vital for improving the efficiency and environmental performance of fresh food logistics.
Related Concept Videos
Reinforcement Schedules
130
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
130
Timing and Consequences on Behavior
78
In operant conditioning, the timing of reinforcement is crucial. For animals like rats and cats, immediate reinforcement (within a few seconds) is much more effective than delayed reinforcement. For example, a food reward for a rat needs to follow within 30 seconds of pressing a bar to be effective.
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
Humans, however, can respond to delayed reinforcers. We often make decisions between immediate small rewards and delayed larger rewards. This ability to delay gratification is a significant...
78

