Semantics through Time: Semi-supervised Segmentation of Aerial Videos with Iterative Label Propagation

Semantic segmentation is a crucial task for robot navigation and safety.\nHowever, current supervised methods require a large amount of pixelwise\nannotations to yield accurate results. Labeling is a tedious and time consuming\nprocess that has hampered progress in low altitude UAV applications. This paper\nmakes an important step towards automatic annotation by introducing SegProp, a\nnovel iterative flow-based method, with a direct connection to spectral\nclustering in space and time, to propagate the semantic labels to frames that\nlack human annotations. The labels are further used in semi-supervised learning\nscenarios. Motivated by the lack of a large video aerial dataset, we also\nintroduce Ruralscapes, a new dataset with high resolution (4K) images and\nmanually-annotated dense labels every 50 frames - the largest of its kind, to\nthe best of our knowledge. Our novel SegProp automatically annotates the\nremaining unlabeled 98% of frames with an accuracy exceeding 90% (F-measure),\nsignificantly outperforming other state-of-the-art label propagation methods.\nMoreover, when integrating other methods as modules inside SegProp's iterative\nlabel propagation loop, we achieve a significant boost over the baseline\nlabels. Finally, we test SegProp in a full semi-supervised setting: we train\nseveral state-of-the-art deep neural networks on the\nSegProp-automatically-labeled training frames and test them on completely novel\nvideos. We convincingly demonstrate, every time, a significant improvement over\nthe supervised scenario.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC