Improving Semantic Segmentation through Spatio-Temporal Consistency Learned from Videos

We leverage unsupervised learning of depth, egomotion, and camera intrinsics\nto improve the performance of single-image semantic segmentation, by enforcing\n3D-geometric and temporal consistency of segmentation masks across video\nframes. The predicted depth, egomotion, and camera intrinsics are used to\nprovide an additional supervision signal to the segmentation model,\nsignificantly enhancing its quality, or, alternatively, reducing the number of\nlabels the segmentation model needs. Our experiments were performed on the\nScanNet dataset.\n

Paper

References (17)

Scroll for more · 5 remaining

Similar papers

© 2026 NYSGPT2525 LLC