Abstract Pursuit-evasion scenarios are typical and significant area of study multi-agent behaviours. While previous research has primarily focused on training pursuit strategies on idealized flat, this paper explores pursuit strategies on heterogeneous ground that rely on ground features from the effects of friction and viscous forces. Since agent interacts with heterogeneous ground and changes its acceleration, we perform multi-agent modelling. Aiming at the multi-agent deep deterministic policy gradient (MADDPG) algorithm’s low training efficiency and slow convergence speed in pursuit learning, we use the prioritized experience selection mechanism of MADDPG algorithm (PES-MADDPG) which improve the experience extraction mechanism based on the policy evaluation function error and the experience extraction training frequency. With this approach, the design of rewards results in faster convergence of reward values. The experimental findings confirm that the suggested method delivers better performance and effectiveness for pursuers and evaders on heterogeneous ground, and both can acquire the relevant movement strategies through training.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex