自动驾驶汽车与路过者互动的决策:一种意识到风险的深度强化学习方法
Ziqian Zhang1, Haojie Li1, Tiantian Chen2
1School of Transportation, Southeast University, China; Jiangsu Key Laboratory of Urban ITS, China; Jiangsu Province Collaborative Innovation Center of Modern Urban Traffic Technologies, China.
一种新的风险意识深度强化学习 (DRL) 方法通过预测路过风险来提高自动驾驶汽车 (AV) 的安全性. 这种方法可以改善复杂的行人与车辆交互中的决策,提高安全性和效率.
科学领域:
- 人工智能的人工智能
- 机器人技术 机器人技术 机器人技术
- 运输工程 运输工程
背景情况:
- 由于反应时间有限,捷径行走对驾驶员和行人构成重大风险.
- 自动驾驶汽车 (AV) 技术为减轻这些风险提供了潜在的解决方案.
- 当前的自动驾驶决策框架可能无法充分解决 jaywalker - 车辆交互的复杂性.
研究的目的:
- 开发和评估一种风险意识的深度强化学习 (DRL) 方法.
- 在涉及路过者的场景中增强AV决策.
- 提高自主导航的安全性和效率.
主要方法:
- 使用编码器-解码器模型的风险预测模块被集成到DRL框架中.
- 该方法在模拟环境 (Anylogic) 中与真实世界的运动数据进行了测试.
- 在不同的情景中评估了AV政策,其中的 jaywalker 量有所不同.
主要成果:
- 建议的风险意识DRL方法在安全指标上显著超过了基线政策,稳定了"低TTC比率"在0.08附近.
- 该方法的优越性随着场景的复杂性 (更高的 jaywalker 卷) 的增加而增加.
- 效率表现排名第二,复杂场景中AV延迟最小.
结论:
- 风险意识的DRL方法使AV能够有效地预测和减轻街行风险.
- 这项技术通过改善高风险地区的AV导航来提高行人安全.
- 拟议的方法为更安全,更高效的自动驾驶提供了实际解决方案.
更多相关视频
06:28A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
Published on: August 26, 2018
07:42An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
Published on: August 2, 2018
相关概念视频
Decision Making
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
