Why Do Self-Supervised Models Transfer? Investigating the Impact of Invariance on Downstream Tasks

Self-supervised learning is a powerful paradigm for representation learning\non unlabelled images. A wealth of effective new methods based on instance\nmatching rely on data-augmentation to drive learning, and these have reached a\nrough agreement on an augmentation scheme that optimises popular recognition\nbenchmarks. However, there is strong reason to suspect that different tasks in\ncomputer vision require features to encode different (in)variances, and\ntherefore likely require different augmentation strategies. In this paper, we\nmeasure the invariances learned by contrastive methods and confirm that they do\nlearn invariance to the augmentations used and further show that this\ninvariance largely transfers to related real-world changes in pose and\nlighting. We show that learned invariances strongly affect downstream task\nperformance and confirm that different downstream tasks benefit from polar\nopposite (in)variances, leading to performance loss when the standard\naugmentation strategy is used. Finally, we demonstrate that a simple fusion of\nrepresentations with complementary invariances ensures wide transferability to\nall the diverse downstream tasks considered.\n

Paper

References (51)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC