Towards Fine-grained Visual Representations by Combining Contrastive Learning with Image Reconstruction and Attention-weighted Pooling

This paper presents Contrastive Reconstruction, ConRec - a self-supervised\nlearning algorithm that obtains image representations by jointly optimizing a\ncontrastive and a self-reconstruction loss. We showcase that state-of-the-art\ncontrastive learning methods (e.g. SimCLR) have shortcomings to capture\nfine-grained visual features in their representations. ConRec extends the\nSimCLR framework by adding (1) a self-reconstruction task and (2) an attention\nmechanism within the contrastive learning task. This is accomplished by\napplying a simple encoder-decoder architecture with two heads. We show that\nboth extensions contribute towards an improved vector representation for images\nwith fine-grained visual features. Combining those concepts, ConRec outperforms\nSimCLR and SimCLR with Attention-Pooling on fine-grained classification\ndatasets.\n

Paper

References (39)

Scroll for more · 27 remaining

Similar papers

© 2026 NYSGPT2525 LLC