Iterative Knowledge Exchange Between Deep Learning and Space-Time Spectral Clustering for Unsupervised Segmentation in Videos

We propose a dual system for unsupervised object segmentation in video, which\nbrings together two modules with complementary properties: a space-time graph\nthat discovers objects in videos and a deep network that learns powerful object\nfeatures. The system uses an iterative knowledge exchange policy. A novel\nspectral space-time clustering process on the graph produces unsupervised\nsegmentation masks passed to the network as pseudo-labels. The net learns to\nsegment in single frames what the graph discovers in video and passes back to\nthe graph strong image-level features that improve its node-level features in\nthe next iteration. Knowledge is exchanged for several cycles until\nconvergence. The graph has one node per each video pixel, but the object\ndiscovery is fast. It uses a novel power iteration algorithm computing the main\nspace-time cluster as the principal eigenvector of a special Feature-Motion\nmatrix without actually computing the matrix. The thorough experimental\nanalysis validates our theoretical claims and proves the effectiveness of the\ncyclical knowledge exchange. We also perform experiments on the supervised\nscenario, incorporating features pretrained with human supervision. We achieve\nstate-of-the-art level on unsupervised and supervised scenarios on four\nchallenging datasets: DAVIS, SegTrack, YouTube-Objects, and DAVSOD.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC