生成自然主义和关键边界场景:一个双层自适应的深度强化学习方法
Junjie Zhou1, Lin Wang2, Qiang Meng3
1State Key Laboratory of Submarine Geoscience, School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai, 200240, China; College of Engineering, Ocean University of China, Qingdao, 266100, China.
Accident; analysis and prevention
|October 8, 2025
概括
本研究引入了双层自适应深度强化学习 (BADRL) 框架,以创建用于测试自动驾驶汽车的现实驾驶场景. 该BADRL方法提高了关键边界场景生成效率15.89%.
科学领域:
- 自主驾驶系统 自主驾驶系统
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 现实世界驾驶对自动驾驶汽车 (AV) 测试提出了复杂的挑战.
- 有限的自然和关键测试场景阻碍了全面的AV性能评估.
- 现有的方法与公正和有效的场景生成作斗争.
研究的目的:
- 提出一个新的框架,用于为AVs生成现实的和多样化的临界边界场景.
- 开发一种人工智能驱动的方法,用于公正的AV性能评估.
- 提高AV测试环境的效率和真实性.
主要方法:
- 开发了一个双层自适应深度强化学习 (BADRL) 框架.
- 人工智能代理人接受了自然驾驶数据的训练,以模仿现实的行为.
- 引入了一个场景复杂性模型,用于实时复杂性评估和动态升级.
- 不同的交通参与者 (车辆,行人,自行车) 被模拟为复杂的互动.
主要成果:
- 该BADRL框架可以在线实时生成自然主义和关键边界场景.
- 该方法有效地模拟了各种交通参与者之间复杂的交互行为.
- 模拟实验证明了BADRL方法在各种驾驶环境中的有效性.
- 与最先进的方法相比,在关键边界场景生成效率上观察到大约15.89%的改善.
结论:
- BADRL框架在为自动驾驶汽车生成高保真测试场景方面取得了重大进展.
- 该方法解决了当前AV测试中的关键局限性,为更强大,更可靠的自动驾驶系统铺平了道路.
- 拟议的方法提高了关键场景生成的效率和现实性,这对于安全的AV部署至关重要.
相关概念视频
Avoidance Learning and Learned Helplessness
2.5K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
2.5K
Observational Learning
832
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
832
Reinforcement
830
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
830
Reinforcement Schedules
458
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
458

