Related Experiment Video
Updated: Jan 14, 2026

A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
Published on: August 26, 2018
Trustworthy navigation with variational policy in deep reinforcement learning.
Karla Bockrath1, Liam Ernst1, Rohaan Nadeem1
1Chester F. Carlson Center for Imaging Science, Rochester Institute of Technology, Rochester, NY, United States.
This study introduces Trust-Nav, a new framework for trustworthy navigation in mobile robots using deep reinforcement learning (DRL). Trust-Nav quantifies uncertainty for safer navigation in unknown and dynamic environments.
Area of Science:
- Robotics
- Artificial Intelligence
- Machine Learning
Background:
- Developing reliable navigation for mobile robots in dynamic environments is challenging.
- Deep Reinforcement Learning (DRL) struggles with uncertainty estimation in real-world applications.
- Autonomous navigation requires robust obstacle avoidance and mapping without prior knowledge.
Purpose of the Study:
- Introduce a novel trustworthy navigation framework, Trust-Nav.
- Quantify uncertainty in robot action, localization, and map representation.
- Enhance the safety and reliability of DRL-based navigation systems.
Main Methods:
- Utilize variational policy learning with Bayesian variational approximation.
- Combine policy-based and value-based learning for action guidance.
- Embed uncertainty in the reward function using Optimal Experimental Design principles.
Main Results:
- Demonstrate superior performance of Trust-Nav in Gazebo simulations.
- Achieve robust autonomous navigation and mapping capabilities.
- Outperform deterministic DRL approaches in noisy and adversarial conditions.
Conclusions:
- Trust-Nav provides safer and more reliable navigation by integrating uncertainty.
- The framework enables mobile robots to recognize and respond to their limitations.
- Represents a step towards deployable, self-aware robotic systems.
Related Concept Videos
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Reinforcement Schedules
Once a behavior is learned,...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Propagation of Uncertainty from Random Error
Woodward–Hoffmann Selection Rules and Microscopic Reversibility

