Related Experiment Video
Updated: Aug 14, 2026

Deep-Learning Based Multi-Joint Synchronous Tracking for Objective Quantification of Hindlimb Locomotor Kinematics in Rats
Published on: April 3, 2026
PPO-GAT-Follow: Graph-Attention Reinforcement Learning for Robust Robot Person Following in Dense Crowds
Xinyu Zhou1, Yongliang Shi2, Songhao Piao1
1Multi-Agent Robot Research Center, Faculty of Computing, Harbin Institute of Technology, Harbin 150001, China.
Abstract:
Robot person following (RPF) in dense crowds requires a mobile robot to maintain an appropriate relative position with respect to a moving target while avoiding surrounding pedestrians and satisfying rear-following and social constraints. This paper proposes PPO-GAT-Follow, an interaction-aware reinforcement learning framework for dense-crowd RPF under geometric visibility loss with available target-relative pose estimates. The follower, target pedestrian, and surrounding pedestrians are represented as graph nodes, and a graph attention encoder models their local interactions. A task-oriented reward mechanism jointly accounts for target maintenance, visibility preservation, collision avoidance, proximity-aware social compliance, rear position maintenance, post-arrival stabilization, and action stability. Experiments are conducted in IR-SIM under fixed-route and random-route settings, with comparisons against MPC, DWA, SFM, and an adapted SARL baseline. In the fixed-route setting with 12 background pedestrians, PPO-GAT-Follow achieves a task success rate of 98.8% and a collision rate of 1.1%, improving task success by 10.9 percentage points over MPC. In the random-route setting at the training density, it achieves 83.1% task success and an SPL of 0.815, outperforming MPC by 18.3 percentage points in task success; at this density, it also surpasses SARL in the main task-level metrics. Zero-shot evaluations across crowd densities, together with structural and reward ablations, reward weight sensitivity analysis, tolerance shift tests, multi-seed training, and stress testing under target pose noise and heterogeneous pedestrian dynamics, further demonstrate the effectiveness and reliability of the proposed framework. Gazebo-based validation also demonstrates system integration feasibility with localization, point cloud-based surrounding pedestrian perception, tracking, and UWB-like target-relative pose input. Nevertheless, visual target identification, re-identification, and perception-level occlusion recovery remain outside the scope of the present validation.
Related Concept Videos
Automatic Processing and Automatic Social Behavior
Observational Learning