Reward Poisoning in Reinforcement Learning: Attacks Against Unknown Learners in Unknown Environments
We study black-box reward poisoning attacks against reinforcement learning\n(RL), in which an adversary aims to manipulate the rewards to mislead a\nsequence of RL agents with unknown algorithms to learn a nefarious policy in an\nenvironment unknown to the adversary a priori. That is, our attack makes\nminimum assumptions on the prior knowledge of the adversary: it has no initial\nknowledge of the environment or the learner, and neither does it observe the\nlearner's internal mechanism except for its performed actions. We design a\nnovel black-box attack, U2, that can provably achieve a near-matching\nperformance to the state-of-the-art white-box attack, demonstrating the\nfeasibility of reward poisoning even in the most challenging black-box setting.\n
Paper
References (63)
Scroll for more · 38 remaining