在深度强化学习中,可靠的导航具有可变的政策
Karla Bockrath1, Liam Ernst1, Rohaan Nadeem1
1Chester F. Carlson Center for Imaging Science, Rochester Institute of Technology, Rochester, NY, United States.
Frontiers in robotics and AI
|October 24, 2025
概括
本研究介绍了Trust-Nav,这是使用深度强化学习 (DRL) 进行移动机器人可靠导航的新框架. Trust-Nav量化了在未知和动态环境中更安全的导航的不确定性.
科学领域:
- 机器人技术 机器人技术 机器人技术
- 人工智能的人工智能
- 机器学习 机器学习
背景情况:
- 在动态环境中为移动机器人开发可靠的导航是具有挑战性的.
- 深度强化学习 (DRL) 在现实应用中扎着不确定性估计.
- 自主导航需要强大的避难障碍和绘制地图,而无需事先的知识.
研究的目的:
- 引入一个新的可信导航框架,Trust-Nav.
- 量化机器人行动,定位和地图表示中的不确定性.
- 提高基于DRL的导航系统的安全性和可靠性.
主要方法:
- 使用贝叶斯式变量方程近似来利用变量方程学习.
- 将基于政策和基于价值的学习结合为行动指导.
- 在奖励函数中嵌入不确定性,使用最佳实验设计原则.
主要成果:
- 在 Gazebo 模拟中展示 Trust-Nav 的卓越性能.
- 实现强大的自主导航和绘图能力.
- 在噪音和对抗条件下优于决定性DRL方法.
结论:
- 通过整合不确定性,Trust-Nav提供了更安全,更可靠的导航.
- 该框架使移动机器人能够识别并应对其局限性.
- 代表着向可部署,自我意识的机器人系统迈出的一步.
相关概念视频
Reinforcement
826
Positive and negative reinforcement are key concepts in operant conditioning, a learning process where the consequences of a behavior affect the likelihood of that behavior being repeated.
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
826
Observational Learning
824
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
824
Reinforcement Schedules
453
Positive reinforcement is a powerful method for teaching new behaviors to both animals and humans. B.F. Skinner demonstrated this with his experiments using rats in a Skinner box. When a rat pressed a lever, it received a food pellet. This immediate reward encouraged the rat to repeat the behavior. This method, where a reward follows every instance of the behavior, is known as continuous reinforcement. It is highly effective for establishing new behaviors quickly.
Once a behavior is learned,...
Once a behavior is learned,...
453
Decision Making: P-value Method
6.8K
The process of hypothesis testing based on the P-value method includes calculating the P- value using the sample data and interpreting it.
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
6.8K
Propagation of Uncertainty from Random Error
1.6K
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
1.6K
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
3.8K
Electrocyclic reactions, cycloadditions, and sigmatropic rearrangements are concerted pericyclic reactions that proceed via a cyclic transition state. These reactions are stereospecific and regioselective. The stereochemistry of the products depends on the symmetry characteristics of the interacting orbitals and the reaction conditions. Accordingly, pericyclic reactions are classified as either symmetry-allowed or symmetry-forbidden. Woodward and Hoffmann presented the selection criteria for...
3.8K


