Random Expert Distillation: Imitation Learning via Expert Policy Support Estimation

We consider the problem of imitation learning from a finite set of expert\ntrajectories, without access to reinforcement signals. The classical approach\nof extracting the expert's reward function via inverse reinforcement learning,\nfollowed by reinforcement learning is indirect and may be computationally\nexpensive. Recent generative adversarial methods based on matching the policy\ndistribution between the expert and the agent could be unstable during\ntraining. We propose a new framework for imitation learning by estimating the\nsupport of the expert policy to compute a fixed reward function, which allows\nus to re-frame imitation learning within the standard reinforcement learning\nsetting. We demonstrate the efficacy of our reward function on both discrete\nand continuous domains, achieving comparable or better performance than the\nstate of the art under different reinforcement learning algorithms.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC