Guided Uncertainty-Aware Policy Optimization: Combining Learning and Model-Based Strategies for Sample-Efficient Policy Learning

Traditional robotic approaches rely on an accurate model of the environment,\na detailed description of how to perform the task, and a robust perception\nsystem to keep track of the current state. On the other hand, reinforcement\nlearning approaches can operate directly from raw sensory inputs with only a\nreward signal to describe the task, but are extremely sample-inefficient and\nbrittle. In this work, we combine the strengths of model-based methods with the\nflexibility of learning-based methods to obtain a general method that is able\nto overcome inaccuracies in the robotics perception/actuation pipeline, while\nrequiring minimal interactions with the environment. This is achieved by\nleveraging uncertainty estimates to divide the space in regions where the given\nmodel-based policy is reliable, and regions where it may have flaws or not be\nwell defined. In these uncertain regions, we show that a locally learned-policy\ncan be used directly with raw sensory inputs. We test our algorithm, Guided\nUncertainty-Aware Policy Optimization (GUAPO), on a real-world robot performing\npeg insertion. Videos are available at https://sites.google.com/view/guapo-rl\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC