Rethinking preventing class-collapsing in metric learning with margin-based losses

Metric learning seeks perceptual embeddings where visually similar instances\nare close and dissimilar instances are apart, but learned representations can\nbe sub-optimal when the distribution of intra-class samples is diverse and\ndistinct sub-clusters are present. Although theoretically with optimal\nassumptions, margin-based losses such as the triplet loss and margin loss have\na diverse family of solutions. We theoretically prove and empirically show that\nunder reasonable noise assumptions, margin-based losses tend to project all\nsamples of a class with various modes onto a single point in the embedding\nspace, resulting in a class collapse that usually renders the space ill-sorted\nfor classification or retrieval. To address this problem, we propose a simple\nmodification to the embedding losses such that each sample selects its nearest\nsame-class counterpart in a batch as the positive element in the tuple. This\nallows for the presence of multiple sub-clusters within each class. The\nadaptation can be integrated into a wide range of metric learning losses. The\nproposed sampling method demonstrates clear benefits on various fine-grained\nimage retrieval datasets over a variety of existing losses; qualitative\nretrieval results show that samples with similar visual patterns are indeed\ncloser in the embedding space.\n

Paper

References (57)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC