Offline-to-Online Reinforcement Learning via Balanced Replay and Pessimistic Q-Ensemble

Recent advance in deep offline reinforcement learning (RL) has made it\npossible to train strong robotic agents from offline datasets. However,\ndepending on the quality of the trained agents and the application being\nconsidered, it is often desirable to fine-tune such agents via further online\ninteractions. In this paper, we observe that state-action distribution shift\nmay lead to severe bootstrap error during fine-tuning, which destroys the good\ninitial policy obtained via offline RL. To address this issue, we first propose\na balanced replay scheme that prioritizes samples encountered online while also\nencouraging the use of near-on-policy samples from the offline dataset.\nFurthermore, we leverage multiple Q-functions trained pessimistically offline,\nthereby preventing overoptimism concerning unfamiliar actions at novel states\nduring the initial training phase. We show that the proposed method improves\nsample-efficiency and final performance of the fine-tuned robotic agents on\nvarious locomotion and manipulation tasks. Our code is available at:\nhttps://github.com/shlee94/Off2OnRL.\n

Paper

References (46)

Scroll for more · 34 remaining

Similar papers

© 2026 NYSGPT2525 LLC