Hierarchical Policy-Gradient Reinforcement Learning for Multi-Agent Shepherding Control of Non-Cohesive Targets
We propose a decentralized reinforcement learning solution for multi-agent shepherding of non-cohesive targets using policy-gradient methods. Our architecture integrates target-selection with target-driving through Proximal Policy Optimization, enabling continuous action spaces and smoother agent trajectories compared to discrete-action approaches. This model-free framework effectively solves the shepherding problem while exhibiting better performance than model-based solutions previously presented in the literature. Experiments demonstrate our method’s effectiveness and scalability with increased target numbers and limited sensing capabilities.