学会投注:使用强化学习来改善拍卖式智能交叉路口的车辆投标
Giacomo Cabri1, Matteo Lugli1, Manuela Montangero1
1Department of Physics, Informatics and Mathematics, University of Modena e Reggio Emilia, 41125 Modena, Italy.
Sensors (Basel, Switzerland)
|February 24, 2024
概括
自动互联汽车使用强化学习来节省在基于拍卖的交通系统中的钱. 这种智能系统显著降低了成本,节省了高达74%的成本,而不会增加旅行时间.
科学领域:
- 智能运输系统 智能运输系统
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 物联网 (IoT) 的普及将导致拥有自动驾驶汽车和智能管理系统的城市.
- 这些系统需要与城市基础设施和车辆互动,以优化城市移动性.
- 资源管理,特别是节省成本,对于自主系统的经济可行性至关重要.
研究的目的:
- 为自动连接汽车 (ACV) 提出一个强化学习模型.
- 为了使ACV能够在基于拍卖的交叉路口管理系统中节省资源,特别是预算.
- 评估在各种交通条件下节省成本和行程时间之间的权衡.
主要方法:
- 开发了一个使用深度Q学习,一种强化学习的模型.
- 训练了多个模型,并对交通条件的变化进行了训练,以确定最佳性能.
- 将拟议模型的性能与现有和随机策略进行了比较.
主要成果:
- 强化学习模型在不同的交通场景中表现出稳健性.
- 与标准投标人相比,实现了显著的预算节约:在重型交通中至少20%和轻型交通中高达74%.
- 节省的成本大约是随机投标策略的三倍.
- 观察到等待时间的最小增加,表明成本和效率之间的有效平衡.
结论:
- 拟议的强化学习模型为自动驾驶车辆导航节省资源提供了一个实际的解决方案.
- 该模型能够在不影响旅行时间的情况下实现大幅降低成本,这表明该模型可用于现实世界的部署.
- 这种方法非常适合未来的智能城市环境,利用基于拍卖的交通管理.
相关概念视频
The Anchoring-and-Adjustment Heuristic
7.2K
In order to make good decisions, we use our knowledge and our reasoning. Often, this knowledge and reasoning is sound and solid. However, sometimes, we are swayed by biases or by others manipulating a situation. For example, let’s say you and three friends wanted to rent a house and had a combined target budget of $1,600. The realtor shows you only very run-down houses for $1,600 and then shows you a very nice house for $2,000. Might you ask each person to pay more in rent to get the...
7.2K
Reinforcement Schedules
147
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
147
Operant Conditioning Intervention
56
Operant conditioning serves as a foundational principle in therapeutic interventions aimed at modifying maladaptive behaviors. Central to this approach is the notion that behaviors, both adaptive and maladaptive, are learned through reinforcement. By analyzing the environmental factors that reinforce problematic behaviors, clinicians can design interventions to weaken these reinforcements and replace maladaptive behaviors with healthier alternatives.
In operant conditioning, behaviors that are...
In operant conditioning, behaviors that are...
56
PD Controller: Design
229
In automotive engineering, car suspension systems often employ Proportional Derivative (PD) controllers to enhance performance. PD controllers are utilized to adjust the damping force in response to road conditions. A controller, acting as an amplifier with a constant gain, demonstrates proportional control, with output directly mirroring input.
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
Designing a continuous-data controller requires selecting and linking components like adders and integrators, which are fundamental in Proportional,...
229


