3D object detection from monocular images is an ill-posed problem due to the\nprojective entanglement of depth and scale. To overcome this ambiguity, we\npresent a novel self-supervised method for textured 3D shape reconstruction and\npose estimation of rigid objects with the help of strong shape priors and 2D\ninstance masks. Our method predicts the 3D location and meshes of each object\nin an image using differentiable rendering and a self-supervised objective\nderived from a pretrained monocular depth estimation network. We use the KITTI\n3D object detection dataset to evaluate the accuracy of the method. Experiments\ndemonstrate that we can effectively use noisy monocular depth and\ndifferentiable rendering as an alternative to expensive 3D ground-truth labels\nor LiDAR information.\n
Paper
References (43)
Scroll for more · 31 remaining