BSUV-Net 2.0: Spatio-Temporal Data Augmentations for Video-Agnostic Supervised Background Subtraction
Background subtraction (BGS) is a fundamental video processing task which is\na key component of many applications. Deep learning-based supervised algorithms\nachieve very good perforamnce in BGS, however, most of these algorithms are\noptimized for either a specific video or a group of videos, and their\nperformance decreases dramatically when applied to unseen videos. Recently,\nseveral papers addressed this problem and proposed video-agnostic supervised\nBGS algorithms. However, nearly all of the data augmentations used in these\nalgorithms are limited to the spatial domain and do not account for temporal\nvariations that naturally occur in video data. In this work, we introduce\nspatio-temporal data augmentations and apply them to one of the leading\nvideo-agnostic BGS algorithms, BSUV-Net. We also introduce a new\ncross-validation training and evaluation strategy for the CDNet-2014 dataset\nthat makes it possible to fairly and easily compare the performance of various\nvideo-agnostic supervised BGS algorithms. Our new model trained using the\nproposed data augmentations, named BSUV-Net 2.0, significantly outperforms\nstate-of-the-art algorithms evaluated on unseen videos of CDNet-2014. We also\nevaluate the cross-dataset generalization capacity of BSUV-Net 2.0 by training\nit solely on CDNet-2014 videos and evaluating its performance on LASIESTA\ndataset. Overall, BSUV-Net 2.0 provides a ~5% improvement in the F-score over\nstate-of-the-art methods on unseen videos of CDNet-2014 and LASIESTA datasets.\nFurthermore, we develop a real-time variant of our model, that we call Fast\nBSUV-Net 2.0, whose performance is close to the state of the art.\n
Paper
References (47)
Scroll for more · 35 remaining