Efficient large-scale image retrieval with deep feature orthogonality and Hybrid-Swin-Transformers
We present an efficient end-to-end pipeline for largescale landmark\nrecognition and retrieval. We show how to combine and enhance concepts from\nrecent research in image retrieval and introduce two architectures especially\nsuited for large-scale landmark identification. A model with deep orthogonal\nfusion of local and global features (DOLG) using an EfficientNet backbone as\nwell as a novel Hybrid-Swin-Transformer is discussed and details how to train\nboth architectures efficiently using a step-wise approach and a sub-center\narcface loss with dynamic margins are provided. Furthermore, we elaborate a\nnovel discriminative re-ranking methodology for image retrieval. The\nsuperiority of our approach was demonstrated by winning the recognition and\nretrieval track of the Google Landmark Competition 2021.\n