COPA: Certifying Robust Policies for Offline Reinforcement Learning against Poisoning Attacks
As reinforcement learning (RL) has achieved near human-level performance in a\nvariety of tasks, its robustness has raised great attention. While a vast body\nof research has explored test-time (evasion) attacks in RL and corresponding\ndefenses, its robustness against training-time (poisoning) attacks remains\nlargely unanswered. In this work, we focus on certifying the robustness of\noffline RL in the presence of poisoning attacks, where a subset of training\ntrajectories could be arbitrarily manipulated. We propose the first\ncertification framework, COPA, to certify the number of poisoning trajectories\nthat can be tolerated regarding different certification criteria. Given the\ncomplex structure of RL, we propose two certification criteria: per-state\naction stability and cumulative reward bound. To further improve the\ncertification, we propose new partition and aggregation protocols to train\nrobust policies. We further prove that some of the proposed certification\nmethods are theoretically tight and some are NP-Complete problems. We leverage\nCOPA to certify three RL environments trained with different algorithms and\nconclude: (1) The proposed robust aggregation protocols such as temporal\naggregation can significantly improve the certifications; (2) Our certification\nfor both per-state action stability and cumulative reward bound are efficient\nand tight; (3) The certification for different training algorithms and\nenvironments are different, implying their intrinsic robustness properties. All\nexperimental results are available at https://copa-leaderboard.github.io.\n
Paper
References (62)
Scroll for more · 38 remaining