Visuomotor Mechanical Search: Learning to Retrieve Target Objects in Clutter

When searching for objects in cluttered environments, it is often necessary\nto perform complex interactions in order to move occluding objects out of the\nway and fully reveal the object of interest and make it graspable. Due to the\ncomplexity of the physics involved and the lack of accurate models of the\nclutter, planning and controlling precise predefined interactions with accurate\noutcome is extremely hard, when not impossible. In problems where accurate\n(forward) models are lacking, Deep Reinforcement Learning (RL) has shown to be\na viable solution to map observations (e.g. images) to good interactions in the\nform of close-loop visuomotor policies. However, Deep RL is sample inefficient\nand fails when applied directly to the problem of unoccluding objects based on\nimages. In this work we present a novel Deep RL procedure that combines i)\nteacher-aided exploration, ii) a critic with privileged information, and iii)\nmid-level representations, resulting in sample efficient and effective learning\nfor the problem of uncovering a target object occluded by a heap of unknown\nobjects. Our experiments show that our approach trains faster and converges to\nmore efficient uncovering solutions than baselines and ablations, and that our\nuncovering policies lead to an average improvement in the graspability of the\ntarget object, facilitating downstream retrieval applications.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC