Learning Safety-Guaranteed, Non-Greedy Control Barrier Functions Using Reinforcement Learning

Spacecraft rendezvous and proximity operations (RPO) present inherent safety risks to high-value assets, making safety guarantees essential for mission success. However, overly conservative control policies developed for safety can reduce mission efficiency. This work proposes a unified two-stage reinforcement learning (RL) framework that addresses the complementary limitations of a traditional safety approach—Input Constrained Control Barrier Functions (ICCBFs)—in safety-critical, fuel-limited spacecraft control. For a certified safe set $\mathcal{S}$, ICCBFs yield a provably invariant inner set $\mathcal{C}^{*} \subseteq \mathcal{S}$ under input bounds, but the resulting per-step Quadratic Program (QP) is greedy and fuel-inefficient inside $\mathcal{C}^{*}$, and states in the residual set $\mathcal{S} \backslash \mathcal{C}^{*}$ that are recoverable are conservatively abandoned. Stage 1 of the proposed framework learns state-dependent class-$\mathcal{K}_{\infty}$ parameters that adapt the ICCBF/CLF decay rates to embed long-horizon cost awareness while maintaining invariance in $\mathcal{C}^{*}$. Stage 2 then learns a residual barrier $h_{\text{RL}}(\boldsymbol{x})$ that provides recoverability for a subset of $\mathcal{S} \backslash \mathcal{C}^{*}$. At run-time, the framework selects the appropriate barrier formulation (Stage 1 or Stage 2) and solves a lightweight QP under zero-order hold. Both stages are trained with PPO using rewards that penalise constraint violations, control effort, and task-specific metrics. The framework is evaluated on three problems: cruise control, spacecraft rendezvous with a rotating target, and a spacecraft inspection task that maximizes observability while respecting keep-in and keep-out zone constraints. The test cases show a reduction in median fuel consumption compared to the ICCBF baselines by 25.00% to 12.00%, and a 7.00% to 8.00% increase in the number of cases that remain inside the safe set $\mathcal{S}$. These results show that RL can embed long-term cost awareness into safety-critical control structures while maintaining the computational efficiency of QPs.

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC