Plannable Approximations to MDP Homomorphisms: Equivariance under Actions

This work exploits action equivariance for representation learning in\nreinforcement learning. Equivariance under actions states that transitions in\nthe input space are mirrored by equivalent transitions in latent space, while\nthe map and transition functions should also commute. We introduce a\ncontrastive loss function that enforces action equivariance on the learned\nrepresentations. We prove that when our loss is zero, we have a homomorphism of\na deterministic Markov Decision Process (MDP). Learning equivariant maps leads\nto structured latent spaces, allowing us to build a model on which we plan\nthrough value iteration. We show experimentally that for deterministic MDPs,\nthe optimal policy in the abstract MDP can be successfully lifted to the\noriginal MDP. Moreover, the approach easily adapts to changes in the goal\nstates. Empirically, we show that in such MDPs, we obtain better\nrepresentations in fewer epochs compared to representation learning approaches\nusing reconstructions, while generalizing better to new goals than model-free\napproaches.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC