Temporally Coherent Embeddings for Self-Supervised Video Representation Learning

This paper presents TCE: Temporally Coherent Embeddings for self-supervised\nvideo representation learning. The proposed method exploits inherent structure\nof unlabeled video data to explicitly enforce temporal coherency in the\nembedding space, rather than indirectly learning it through ranking or\npredictive proxy tasks. In the same way that high-level visual information in\nthe world changes smoothly, we believe that nearby frames in learned\nrepresentations will benefit from demonstrating similar properties. Using this\nassumption, we train our TCE model to encode videos such that adjacent frames\nexist close to each other and videos are separated from one another. Using TCE\nwe learn robust representations from large quantities of unlabeled video data.\nWe thoroughly analyse and evaluate our self-supervised learned TCE models on a\ndownstream task of video action recognition using multiple challenging\nbenchmarks (Kinetics400, UCF101, HMDB51). With a simple but effective 2D-CNN\nbackbone and only RGB stream inputs, TCE pre-trained representations outperform\nall previous selfsupervised 2D-CNN and 3D-CNN pre-trained on UCF101. The code\nand pre-trained models for this paper can be downloaded at:\nhttps://github.com/csiro-robotics/TCE\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC