Towards Hyperparameter-free Policy Selection for Offline Reinforcement Learning

How to select between policies and value functions produced by different\ntraining algorithms in offline reinforcement learning (RL) -- which is crucial\nfor hyperpa-rameter tuning -- is an important open question. Existing\napproaches based on off-policy evaluation (OPE) often require additional\nfunction approximation and hence hyperparameters, creating a chicken-and-egg\nsituation. In this paper, we design hyperparameter-free algorithms for policy\nselection based on BVFT [XJ21], a recent theoretical advance in value-function\nselection, and demonstrate their effectiveness in discrete-action benchmarks\nsuch as Atari. To address performance degradation due to poor critics in\ncontinuous-action domains, we further combine BVFT with OPE to get the best of\nboth worlds, and obtain a hyperparameter-tuning method for Q-function based OPE\nwith theoretical guarantees as a side product.\n

Paper

References (39)

Scroll for more · 27 remaining

Similar papers

© 2026 NYSGPT2525 LLC