SeqNet: Learning Descriptors for Sequence-based Hierarchical Place Recognition

Visual Place Recognition (VPR) is the task of matching current visual imagery\nfrom a camera to images stored in a reference map of the environment. While\ninitial VPR systems used simple direct image methods or hand-crafted visual\nfeatures, recent work has focused on learning more powerful visual features and\nfurther improving performance through either some form of sequential matcher /\nfilter or a hierarchical matching process. In both cases the performance of the\ninitial single-image based system is still far from perfect, putting\nsignificant pressure on the sequence matching or (in the case of hierarchical\nsystems) pose refinement stages. In this paper we present a novel hybrid system\nthat creates a high performance initial match hypothesis generator using short\nlearnt sequential descriptors, which enable selective control sequential score\naggregation using single image learnt descriptors. Sequential descriptors are\ngenerated using a temporal convolutional network dubbed SeqNet, encoding short\nimage sequences using 1-D convolutions, which are then matched against the\ncorresponding temporal descriptors from the reference dataset to provide an\nordered list of place match hypotheses. We then perform selective sequential\nscore aggregation using shortlisted single image learnt descriptors from a\nseparate pipeline to produce an overall place match hypothesis. Comprehensive\nexperiments on challenging benchmark datasets demonstrate the proposed method\noutperforming recent state-of-the-art methods using the same amount of\nsequential information. Source code and supplementary material can be found at\nhttps://github.com/oravus/seqNet.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC