We leverage unsupervised learning of depth, egomotion, and camera intrinsics\nto improve the performance of single-image semantic segmentation, by enforcing\n3D-geometric and temporal consistency of segmentation masks across video\nframes. The predicted depth, egomotion, and camera intrinsics are used to\nprovide an additional supervision signal to the segmentation model,\nsignificantly enhancing its quality, or, alternatively, reducing the number of\nlabels the segmentation model needs. Our experiments were performed on the\nScanNet dataset.\n
Paper
References (17)
Scroll for more · 5 remaining