Sparse Auxiliary Networks for Unified Monocular Depth Prediction and Completion

Estimating scene geometry from data obtained with cost-effective sensors is\nkey for robots and self-driving cars. In this paper, we study the problem of\npredicting dense depth from a single RGB image (monodepth) with optional sparse\nmeasurements from low-cost active depth sensors. We introduce Sparse Auxiliary\nNetworks (SANs), a new module enabling monodepth networks to perform both the\ntasks of depth prediction and completion, depending on whether only RGB images\nor also sparse point clouds are available at inference time. First, we decouple\nthe image and depth map encoding stages using sparse convolutions to process\nonly the valid depth map pixels. Second, we inject this information, when\navailable, into the skip connections of the depth prediction network,\naugmenting its features. Through extensive experimental analysis on one indoor\n(NYUv2) and two outdoor (KITTI and DDAD) benchmarks, we demonstrate that our\nproposed SAN architecture is able to simultaneously learn both tasks, while\nachieving a new state of the art in depth prediction by a significant margin.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC