Unsupervised Video Representation Learning by Bidirectional Feature Prediction

This paper introduces a novel method for self-supervised video representation\nlearning via feature prediction. In contrast to the previous methods that focus\non future feature prediction, we argue that a supervisory signal arising from\nunobserved past frames is complementary to one that originates from the future\nframes. The rationale behind our method is to encourage the network to explore\nthe temporal structure of videos by distinguishing between future and past\ngiven present observations. We train our model in a contrastive learning\nframework, where joint encoding of future and past provides us with a\ncomprehensive set of temporal hard negatives via swapping. We empirically show\nthat utilizing both signals enriches the learned representations for the\ndownstream task of action recognition. It outperforms independent prediction of\nfuture and past.\n

Paper

References (45)

Scroll for more · 33 remaining

Similar papers

© 2026 NYSGPT2525 LLC