Related Experiment Video
Updated: Jun 13, 2026

A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
Published on: August 26, 2018
Research on Reinforcement Learning-Based Autonomous Navigation and Obstacle Avoidance Methods for AGVs in Unknown
Tianye Luo1, Jing Hu1, Bangcheng Zhang2
1School of Mechatronical Engineering, Changchun University of Science and Technology, Changchun 130022, China.
Abstract:
Reinforcement learning (RL) represents an effective approach for developing autonomous navigation and obstacle avoidance capabilities in hospital automated guided vehicles (AGVs). However, real-world adoption is challenged by the need for carefully designed reward functions, low sample efficiency, and slow convergence behaviour. To effectively address these issues, in this work, BEAGM-PPO, a reinforcement learning framework tailored for unknown hospital environments, was proposed. A reference model was initially employed to improve sample efficiency by directing the agent's learning process. The reference model consists of expert demonstrations and policy derivation mechanisms. During the expert demonstration phase, human experts perform the required tasks and generate state-action pair datasets for training. During the policy derivation phase, demonstration data, behaviour cloning, and uncertainty estimation were used to derive the imitated expert policy. An ant colony optimization (ACO)-inspired pheromone mechanism and a memory replay strategy were incorporated to improve target-oriented action selection and supress unnecessary exploration. Experiments conducted in typical 3D simulation scenarios demonstrated that the proposed method achieved the highest arrival rate compared with baseline models. Moreover, the integrated imitation learning approach enables uncertainty estimation for both the policy and the model, while expanded training datasets further enhance performance. Overall, the results prove that BEAGM-PPO serves as a solid theoretical foundation for autonomous navigation in hospital AGVs.
