The feasibility theory of constrained reinforcement learning: a tutorial study

Satisfying safety constraints is a priority concern when solving optimal control problems (OCPs). Due to the existence of infeasibility phenomenon, where a constraint-satisfying solution cannot be found, it is necessary to identify a feasible region before implementing a policy. Existing feasibility theories in model predictive control (MPC) only apply to the case where the policy is a solution to a specific OCP, either optimal or suboptimal. Such feasibility is essentially feasibility of OCPs, not that of policies or states. However, reinforcement learning (RL), as another important control method, represents a policy as a mapping from state to action, which itself is decoupled with OCPs. An RL policy may not be a solution to a specific OCP, that is, not satisfying the OCP’s constraints, especially at an early stage of training. Feasibility analysis of such inadequately trained policies is necessary for safety improvement in RL, but that is not available under existing MPC feasibility theories. This paper proposes a feasibility theory that applies to both MPC and RL by decoupling states, constraints and policies. Starting from a state, different constraints can be constructed, and different policies can be applied. This study reveals that feasibility depends on the combination of these three elements, extending the traditional viewpoint of OCP-specific feasibility theories in MPC. The basis of our theory is to distinguish initial and endless, state and policy feasibility, and their corresponding feasible regions. Based on these concepts, this study analyzes the containment relationships between different feasible regions, which enables to describe feasibility under arbitrary combinations of states, constraints and policies. This study further provides virtual-time constraint design rules along with a practical design tool called feasibility function that helps to achieve the maximum feasible region. The feasibility function either represents a control invariant set or aggregates infinite steps of constraints into a single one. This study reviews most of existing constraint formulations and points out that they are essentially applications of feasibility functions in different forms. This study demonstrates the feasibility theory by visualizing different feasible regions under both MPC and RL policies in an emergency braking control task.

Paper

Similar papers

© 2026 NYSGPT2525 LLC