Indirect Object-to-Robot Pose Estimation from an External Monocular RGB Camera

We present a robotic grasping system that uses a single external monocular\nRGB camera as input. The object-to-robot pose is computed indirectly by\ncombining the output of two neural networks: one that estimates the\nobject-to-camera pose, and another that estimates the robot-to-camera pose.\nBoth networks are trained entirely on synthetic data, relying on domain\nrandomization to bridge the sim-to-real gap. Because the latter network\nperforms online camera calibration, the camera can be moved freely during\nexecution without affecting the quality of the grasp. Experimental results\nanalyze the effect of camera placement, image resolution, and pose refinement\nin the context of grasping several household objects. We also present results\non a new set of 28 textured household toy grocery objects, which have been\nselected to be accessible to other researchers. To aid reproducibility of the\nresearch, we offer 3D scanned textured models, along with pre-trained weights\nfor pose estimation.\n

Paper

References (38)

Scroll for more · 26 remaining

Similar papers

© 2026 NYSGPT2525 LLC