FP3O: Enabling Proximal Policy Optimization in Multi-Agent Cooperation with Parameter-Sharing Versatility

Existing multiagent proximal policy optimization (PPO) algorithms come at the cost of limited generalizability on different parameter-sharing configurations [e.g., full, partial, and nonparameter sharing (NoPS)] when extending the theoretical guarantee of PPO to cooperative multiagent reinforcement learning (MARL). In this study, we introduce a general-purpose method to address this challenge for multiagent PPO, enabling it with parameter-sharing versatility. Our proposal, the full-pipeline paradigm, employs various equivalent decompositions of the advantage function to establish multiple parallel optimization pipelines. We provide theoretical analysis for this procedure: it ensures consistent policy improvement across different types of parameter sharing, establishing a theoretically and practically aligned guarantee. To instantiate this process, we develop a practical algorithm termed full-pipeline PPO (FP3O). We empirically validate the effectiveness and versatility of FP3O through extensive evaluations on various multiagent tasks. Our results demonstrate that FP3O not only outperforms other baselines but also exhibits remarkable versatility in all common parameter-sharing standards.

Paper

Similar papers

© 2026 NYSGPT2525 LLC