Self-supervised video representation methods typically focus on the\nrepresentation of temporal attributes in videos. However, the role of\nstationary versus non-stationary attributes is less explored: Stationary\nfeatures, which remain similar throughout the video, enable the prediction of\nvideo-level action classes. Non-stationary features, which represent temporally\nvarying attributes, are more beneficial for downstream tasks involving more\nfine-grained temporal understanding, such as action segmentation. We argue that\na single representation to capture both types of features is sub-optimal, and\npropose to decompose the representation space into stationary and\nnon-stationary features via contrastive learning from long and short views,\ni.e. long video sequences and their shorter sub-sequences. Stationary features\nare shared between the short and long views, while non-stationary features\naggregate the short views to match the corresponding long view. To empirically\nverify our approach, we demonstrate that our stationary features work\nparticularly well on an action recognition downstream task, while our\nnon-stationary features perform better on action segmentation. Furthermore, we\nanalyse the learned representations and find that stationary features capture\nmore temporally stable, static attributes, while non-stationary features\nencompass more temporally varying ones.\n
Paper
References (47)
Scroll for more · 35 remaining