City-Scale Visual Place Recognition with Deep Local Features Based on Multi-Scale Ordered VLAD Pooling
Visual place recognition is the task of recognizing a place depicted in an\nimage based on its pure visual appearance without metadata. In visual place\nrecognition, the challenges lie upon not only the changes in lighting\nconditions, camera viewpoint, and scale but also the characteristic of\nscene-level images and the distinct features of the area. To resolve these\nchallenges, one must consider both the local discriminativeness and the global\nsemantic context of images. On the other hand, the diversity of the datasets is\nalso particularly important to develop more general models and advance the\nprogress of the field. In this paper, we present a fully-automated system for\nplace recognition at a city-scale based on content-based image retrieval. Our\nmain contributions to the community lie in three aspects. Firstly, we take a\ncomprehensive analysis of visual place recognition and sketch out the unique\nchallenges of the task compared to general image retrieval tasks. Next, we\npropose yet a simple pooling approach on top of convolutional neural network\nactivations to embed the spatial information into the image representation\nvector. Finally, we introduce new datasets for place recognition, which are\nparticularly essential for application-based research. Furthermore, throughout\nextensive experiments, various issues in both image retrieval and place\nrecognition are analyzed and discussed to give some insights into improving the\nperformance of retrieval models in reality.\n The dataset used in this paper can be found at\nhttps://github.com/canhld94/Daejeon520\n