PPO-GAT-Follow: Graph-Attention Reinforcement Learning for Robust Robot Person Following in Dense Crowds

Robot person following (RPF) in dense crowds requires a mobile robot to maintain an appropriate relative position with respect to a moving target while avoiding surrounding pedestrians and satisfying rear-following and social constraints. This paper proposes PPO-GAT-Follow, an interaction-aware reinforcement learning framework for dense-crowd RPF under geometric visibility loss with available target-relative pose estimates. The follower, target pedestrian, and surrounding pedestrians are represented as graph nodes, and a graph attention encoder models their local interactions. A task-oriented reward mechanism jointly accounts for target maintenance, visibility preservation, collision avoidance, proximity-aware social compliance, rear position maintenance, post-arrival stabilization, and action stability. Experiments are conducted in IR-SIM under fixed-route and random-route settings, with comparisons against MPC, DWA, SFM, and an adapted SARL baseline. In the fixed-route setting with 12 background pedestrians, PPO-GAT-Follow achieves a task success rate of 98.8% and a collision rate of 1.1%, improving task success by 10.9 percentage points over MPC. In the random-route setting at the training density, it achieves 83.1% task success and an SPL of 0.815, outperforming MPC by 18.3 percentage points in task success; at this density, it also surpasses SARL in the main task-level metrics. Zero-shot evaluations across crowd densities, together with structural and reward ablations, reward weight sensitivity analysis, tolerance shift tests, multi-seed training, and stress testing under target pose noise and heterogeneous pedestrian dynamics, further demonstrate the effectiveness and reliability of the proposed framework. Gazebo-based validation also demonstrates system integration feasibility with localization, point cloud-based surrounding pedestrian perception, tracking, and UWB-like target-relative pose input. Nevertheless, visual target identification, re-identification, and perception-level occlusion recovery remain outside the scope of the present validation.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC