Monocular Differentiable Rendering for Self-Supervised 3D Object Detection

3D object detection from monocular images is an ill-posed problem due to the\nprojective entanglement of depth and scale. To overcome this ambiguity, we\npresent a novel self-supervised method for textured 3D shape reconstruction and\npose estimation of rigid objects with the help of strong shape priors and 2D\ninstance masks. Our method predicts the 3D location and meshes of each object\nin an image using differentiable rendering and a self-supervised objective\nderived from a pretrained monocular depth estimation network. We use the KITTI\n3D object detection dataset to evaluate the accuracy of the method. Experiments\ndemonstrate that we can effectively use noisy monocular depth and\ndifferentiable rendering as an alternative to expensive 3D ground-truth labels\nor LiDAR information.\n

Paper

References (43)

Scroll for more · 31 remaining

Similar papers

© 2026 NYSGPT2525 LLC