Intermediate Layers Matter in Momentum Contrastive Self Supervised Learning

We show that bringing intermediate layers' representations of two augmented\nversions of an image closer together in self-supervised learning helps to\nimprove the momentum contrastive (MoCo) method. To this end, in addition to the\ncontrastive loss, we minimize the mean squared error between the intermediate\nlayer representations or make their cross-correlation matrix closer to an\nidentity matrix. Both loss objectives either outperform standard MoCo, or\nachieve similar performances on three diverse medical imaging datasets:\nNIH-Chest Xrays, Breast Cancer Histopathology, and Diabetic Retinopathy. The\ngains of the improved MoCo are especially large in a low-labeled data regime\n(e.g. 1% labeled data) with an average gain of 5% across three datasets. We\nanalyze the models trained using our novel approach via feature similarity\nanalysis and layer-wise probing. Our analysis reveals that models trained via\nour approach have higher feature reuse compared to a standard MoCo and learn\ninformative features earlier in the network. Finally, by comparing the output\nprobability distribution of models fine-tuned on small versus large labeled\ndata, we conclude that our proposed method of pre-training leads to lower\nKolmogorov-Smirnov distance, as compared to a standard MoCo. This provides\nadditional evidence that our proposed method learns more informative features\nin the pre-training phase which could be leveraged in a low-labeled data\nregime.\n

Paper

References (36)

Scroll for more · 24 remaining

Similar papers

© 2026 NYSGPT2525 LLC