Learning Barrier Certificates: Towards Safe Reinforcement Learning with Zero Training-time Violations

Training-time safety violations have been a major concern when we deploy\nreinforcement learning algorithms in the real world. This paper explores the\npossibility of safe RL algorithms with zero training-time safety violations in\nthe challenging setting where we are only given a safe but trivial-reward\ninitial policy without any prior knowledge of the dynamics model and additional\noffline data. We propose an algorithm, Co-trained Barrier Certificate for Safe\nRL (CRABS), which iteratively learns barrier certificates, dynamics models, and\npolicies. The barrier certificates, learned via adversarial training, ensure\nthe policy's safety assuming calibrated learned dynamics model. We also add a\nregularization term to encourage larger certified regions to enable better\nexploration. Empirical simulations show that zero safety violations are already\nchallenging for a suite of simple environments with only 2-4 dimensional state\nspace, especially if high-reward policies have to visit regions near the safety\nboundary. Prior methods require hundreds of violations to achieve decent\nrewards on these tasks, whereas our proposed algorithms incur zero violations.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC