6-DoF Pose Estimation of Household Objects for Robotic Manipulation: An Accessible Dataset and Benchmark

We present a new dataset for 6-DoF pose estimation of known objects, with a\nfocus on robotic manipulation research. We propose a set of toy grocery\nobjects, whose physical instantiations are readily available for purchase and\nare appropriately sized for robotic grasping and manipulation. We provide 3D\nscanned textured models of these objects, suitable for generating synthetic\ntraining data, as well as RGBD images of the objects in challenging, cluttered\nscenes exhibiting partial occlusion, extreme lighting variations, multiple\ninstances per image, and a large variety of poses. Using semi-automated\nRGBD-to-model texture correspondences, the images are annotated with ground\ntruth poses accurate within a few millimeters. We also propose a new pose\nevaluation metric called ADD-H based on the Hungarian assignment algorithm that\nis robust to symmetries in object geometry without requiring their explicit\nenumeration. We share pre-trained pose estimators for all the toy grocery\nobjects, along with their baseline performance on both validation and test\nsets. We offer this dataset to the community to help connect the efforts of\ncomputer vision researchers with the needs of roboticists.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC