基于深度强化学习的端到端自主导航,具有生存惩罚函数
Shyr-Long Jeng1, Chienhsun Chiang2
1Department of Mechanical Engineering, Lunghwa University of Science and Technology, Taoyuan City 333326, Taiwan.
Sensors (Basel, Switzerland)
|October 28, 2023
概括
本研究引入了深度强化学习方法,用于在未知的动态环境中实现自主导航. 该方法通过使用新的奖励功能来增强机器人的生存和目标实现,从而实现无碰撞路径规划.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 在动态的,没有地图的环境中,自主导航带来了重大挑战.
- 传统的方法与稀疏的奖励和复杂的障碍回避作斗争.
研究的目的:
- 为自主导航提出一个端到端的深度强化学习 (DRL) 方法.
- 为了使非全方位的轮式移动机器人 (WMR) 能够有效地在动态的,没有地图的环境中导航.
- 为了解决稀疏奖励问题,并确保无碰撞的路径规划.
主要方法:
- 利用了两个行为者-关键 (AC) 框架:深度决定性政策梯度 (DDPG) 和双延迟的 DDPG (TD3).
- 引入了一个全面的奖励功能,包括一个生存惩罚,以引导WMR实现目标.
- 连接连续的事件,以增加障碍场景的累积罚款,防止训练失败.
主要成果:
- 在各种场景 (无障碍,停车场,十字路口,多重障碍) 中进行的模拟证明了该方法的效率和安全性.
- 与DDPG相比,TD3算法在训练过程中显示出更快的融合和更大的稳定性.
- 在评估过程中,TD3实现了更高的任务执行成功率.
结论:
- 拟议的DRL方法具有生存惩罚功能,可以有效地在具有挑战性的环境中实现自主导航.
- 与DDPG相比,TD3算法在训练效率和导航成功率方面提供了卓越的性能.
- 该方法为WMR无碰撞路径规划提供了强大的解决方案.
相关概念视频
Avoidance Learning and Learned Helplessness
1.7K
Avoidance learning and learned helplessness are critical concepts in understanding behavioral responses to negative stimuli.
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
1.7K
Rolling Resistance: Problem Solving
352
Rolling resistance, also known as rolling friction, is the force that resists the motion of a rolling object, such as a wheel, tire, or ball, when it moves over a surface. It is caused by the deformation of the object and the surface in contact with each other, as well as other factors like internal friction, hysteresis, and energy losses within the materials. Rolling resistance opposes the object's motion, requiring additional energy to overcome it and maintain movement. In practical...
352
Reinforcement
221
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
221
Survival Tree
88
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
88
Hydraulic Jump: Problem Solving
65
To analyze a hydraulic jump in a rectangular channel with a flow speed of 6 meters per second, follow these steps:Calculate Effective Upstream Velocity:When the downstream gate closes, a hydraulic jump forms, traveling upstream at 2 meters per second. This wave speed combines with the initial channel flow velocity, creating an effective upstream velocity.Identify Flow Velocities Before and After the Hydraulic Jump:Upstream of the hydraulic jump, the effective flow velocity includes both the...
65
Observational Learning
188
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
188


