Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category Reconstruction

Traditional approaches for learning 3D object categories have been\npredominantly trained and evaluated on synthetic datasets due to the\nunavailability of real 3D-annotated category-centric data. Our main goal is to\nfacilitate advances in this field by collecting real-world data in a magnitude\nsimilar to the existing synthetic counterparts. The principal contribution of\nthis work is thus a large-scale dataset, called Common Objects in 3D, with real\nmulti-view images of object categories annotated with camera poses and ground\ntruth 3D point clouds. The dataset contains a total of 1.5 million frames from\nnearly 19,000 videos capturing objects from 50 MS-COCO categories and, as such,\nit is significantly larger than alternatives both in terms of the number of\ncategories and objects. We exploit this new dataset to conduct one of the first\nlarge-scale "in-the-wild" evaluations of several new-view-synthesis and\ncategory-centric 3D reconstruction methods. Finally, we contribute NerFormer -\na novel neural rendering method that leverages the powerful Transformer to\nreconstruct an object given a small number of its views. The CO3D dataset is\navailable at https://github.com/facebookresearch/co3d .\n

Paper

References (74)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC