A Hierarchical Dual Model of Environment- and Place-Specific Utility for Visual Place Recognition
Visual Place Recognition (VPR) approaches have typically attempted to match\nplaces by identifying visual cues, image regions or landmarks that have high\n``utility'' in identifying a specific place. But this concept of utility is not\nsingular - rather it can take a range of forms. In this paper, we present a\nnovel approach to deduce two key types of utility for VPR: the utility of\nvisual cues `specific' to an environment, and to a particular place. We employ\ncontrastive learning principles to estimate both the environment- and\nplace-specific utility of Vector of Locally Aggregated Descriptors (VLAD)\nclusters in an unsupervised manner, which is then used to guide local feature\nmatching through keypoint selection. By combining these two utility measures,\nour approach achieves state-of-the-art performance on three challenging\nbenchmark datasets, while simultaneously reducing the required storage and\ncompute time. We provide further analysis demonstrating that unsupervised\ncluster selection results in semantically meaningful results, that finer\ngrained categorization often has higher utility for VPR than high level\nsemantic categorization (e.g. building, road), and characterise how these two\nutility measures vary across different places and environments. Source code is\nmade publicly available at https://github.com/Nik-V9/HEAPUtil.\n