Visual place recognition is a challenging task in computer vision and a key\ncomponent of camera-based localization and navigation systems. Recently,\nConvolutional Neural Networks (CNNs) achieved high results and good\ngeneralization capabilities. They are usually trained using pairs or triplets\nof images labeled as either similar or dissimilar, in a binary fashion. In\npractice, the similarity between two images is not binary, but continuous.\nFurthermore, training these CNNs is computationally complex and involves costly\npair and triplet mining strategies. We propose a Generalized Contrastive loss\n(GCL) function that relies on image similarity as a continuous measure, and use\nit to train a siamese CNN. Furthermore, we present three techniques for\nautomatic annotation of image pairs with labels indicating their degree of\nsimilarity, and deploy them to re-annotate the MSLS, TB-Places, and 7Scenes\ndatasets. We demonstrate that siamese CNNs trained using the GCL function and\nthe improved annotations consistently outperform their binary counterparts. Our\nmodels trained on MSLS outperform the state-of-the-art methods, including\nNetVLAD, NetVLAD-SARE, AP-GeM and Patch-NetVLAD, and generalize well on the\nPittsburgh30k, Tokyo 24/7, RobotCar Seasons v2 and Extended CMU Seasons\ndatasets. Furthermore, training a siamese network using the GCL function does\nnot require complex pair mining. We release the source code at\nhttps://github.com/marialeyvallina/generalized_contrastive_loss.\n
Paper
References (76)
Scroll for more · 38 remaining