"What's This?" -- Learning to Segment Unknown Objects from Manipulation Sequences

We present a novel framework for self-supervised grasped object segmentation\nwith a robotic manipulator. Our method successively learns an agnostic\nforeground segmentation followed by a distinction between manipulator and\nobject solely by observing the motion between consecutive RGB frames. In\ncontrast to previous approaches, we propose a single, end-to-end trainable\narchitecture which jointly incorporates motion cues and semantic knowledge.\nFurthermore, while the motion of the manipulator and the object are substantial\ncues for our algorithm, we present means to robustly deal with distraction\nobjects moving in the background, as well as with completely static scenes. Our\nmethod neither depends on any visual registration of a kinematic robot or 3D\nobject models, nor on precise hand-eye calibration or any additional sensor\ndata. By extensive experimental evaluation we demonstrate the superiority of\nour framework and provide detailed insights on its capability of dealing with\nthe aforementioned extreme cases of motion. We also show that training a\nsemantic segmentation network with the automatically labeled data achieves\nresults on par with manually annotated training data. Code and pretrained model\nare available at https://github.com/DLR-RM/DistinctNet.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC