Learning Monocular Depth in Dynamic Scenes via Instance-Aware Projection Consistency

We present an end-to-end joint training framework that explicitly models\n6-DoF motion of multiple dynamic objects, ego-motion and depth in a monocular\ncamera setup without supervision. Our technical contributions are three-fold.\nFirst, we highlight the fundamental difference between inverse and forward\nprojection while modeling the individual motion of each rigid object, and\npropose a geometrically correct projection pipeline using a neural forward\nprojection module. Second, we design a unified instance-aware photometric and\ngeometric consistency loss that holistically imposes self-supervisory signals\nfor every background and object region. Lastly, we introduce a general-purpose\nauto-annotation scheme using any off-the-shelf instance segmentation and\noptical flow models to produce video instance segmentation maps that will be\nutilized as input to our training pipeline. These proposed elements are\nvalidated in a detailed ablation study. Through extensive experiments conducted\non the KITTI and Cityscapes dataset, our framework is shown to outperform the\nstate-of-the-art depth and motion estimation methods. Our code, dataset, and\nmodels are available at https://github.com/SeokjuLee/Insta-DM .\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC