Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning

We study a security threat to reinforcement learning where an attacker\npoisons the learning environment to force the agent into executing a target\npolicy chosen by the attacker. As a victim, we consider RL agents whose\nobjective is to find a policy that maximizes average reward in undiscounted\ninfinite-horizon problem settings. The attacker can manipulate the rewards or\nthe transition dynamics in the learning environment at training-time and is\ninterested in doing so in a stealthy manner. We propose an optimization\nframework for finding an \\emph{optimal stealthy attack} for different measures\nof attack cost. We provide sufficient technical conditions under which the\nattack is feasible and provide lower/upper bounds on the attack cost. We\ninstantiate our attacks in two settings: (i) an \\emph{offline} setting where\nthe agent is doing planning in the poisoned environment, and (ii) an\n\\emph{online} setting where the agent is learning a policy using a\nregret-minimization framework with poisoned feedback. Our results show that the\nattacker can easily succeed in teaching any target policy to the victim under\nmild conditions and highlight a significant security threat to reinforcement\nlearning agents in practice.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC