Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance

Safe reinforcement learning (RL) typically asks what an agent should do. We ask when it needs to act, and show that a single policy can jointly learn control inputs and communication-efficient timing decisions under a pointwise Lyapunov safety shield. We scope the framework to stabilization around a known equilibrium, where CARE-based LQR backups, Lyapunov certificates, and classical Lyapunov- STC are well defined, enabling a clean comparison against the analytical baseline. A run-time assurance (RTA) layer overrides the policy pointwise via a one-step-ahead Lyapunov prediction and a precomputed LQR backup, providing a strictly stronger guarantee than constrained MDP methods that enforce safety only in expectation. On an inverted pendulum, cart–pole, and planar quadrotor, the learned policy achieves 1.91×, 1.45×, and 3.51× higher mean inter-sample interval (MSI) than a classical Lyapunov-triggered baseline; a fixed LQR controller at the same average rate is unstable on all three plants, showing that adaptive timing, not a lower average rate, is what makes sparsity safe. A CARE-derived Lyapunov reward transfers across environments without redesign, with a single weight wc controlling the stability–communication tradeoff; ablations confirm the RTA shield is essential, with its removal reducing MSI by 1.27–1.84× and degrading state norms. A preference-conditioned extension recovers the full tradeoff frontier from a single model at 2 11 of training compute, and SAC experiments confirm the results are algorithm-agnostic across discrete and continuous domains. A 12-state 3D quadrotor case study extends the framework to higher-dimensional systems where classical STC design is analytically intractable: a SAC agent reaches MSI = 0.302 s (94% of τmax) at 0% RTA, while classical Lyapunov-STC remains pinned at τmin and a fixed-rate LQR controller at the same average interval crashes within two control updates. Robustness to ±30% plant-mass variation and additive disturbances confirms graceful degradation, with the RTA absorbing what the learned policy cannot.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC