Image retrieval survey
Image retrieval aims to find images from a database that are relevant to a given query. This crucially differs from clustering in that clustering requires both finding the clusters and assigning the images to them; image retrieval techniques are very relevant to the sub-task of cluster assignment but not to the sub-task of finding the clusters.
The fundamental approach in image retrieval is to assess the similarity among image features. Current approaches focus on two kinds of image representations: global features and local features. For global representations, [1, 2, 3, 4, 5]extracts activations from deep CNNs and aggregates them for obtaining global features. For local representations, [6, 7, 8, 9, 10, 11, 12] proposed well-embedded representations for all regions of interest. Recent state-of-the-art methods [7, 13, 4, 14, 15] typically followed a two-stage paradigm: initially, candidates are retrieved using global features, and then they are re-ranked with local features. Recently, [16, 17, 18, 19] proposed to condition retrieval on user-specified language.
[1] A. Babenko, A. Slesarev, A. Chigorin, and V. Lempitsky. Neural codes for image retrieval. European
Conference on Computer Vision, 2014.
[2] G. Tolias, R. Sicre, and H. Jegou. Particular object retrieval with integral max-pooling of cnn ´
activations. arXiv: 1511.05879, 2015.
[3] A. Gordo, J. Almazan, J. Revaud, and D. Larlus. Deep image retrieval: Learning global representa- ´
tions for image search. European Conference on Computer Vision, 2016.
[4] B. Cao, A. Araujo, and J. Sim. Unifying deep local and global features for image search. European
Conference on Computer Vision, 2020a.
[5] S. Lee, S. Lee, H. Seong, and E. Kim. Revisiting self-similarity: Structural embedding for image
retrieval. Conference on Computer Vision and Pattern Recognition, 2023.
[6] K. M. Yi, E. Trulls, V. Lepetit, and P. Fua. Lift: Learned invariant feature transform. European
Conference on Computer Vision, pages 467–483, 2016.
[7] H. Noh, A. Araujo, J. Sim, T. Weyand, and B. Han. Large-scale image retrieval with attentive deep
local features. International Conference on Computer Vision, 2017.
[8] D. P. Vassileios Balntas, Edgar Riba and K. Mikolajczyk. Learning local feature descriptors with
triplets and shallow convolutional neural networks. Proceedings of the British Machine Vision
Conference (BMVC), 2016.
[9] D. DeTone, T. Malisiewicz, and A. Rabinovich. Superpoint: Self-supervised interest point detection
and description. Computer Vision and Pattern Recognition Workshops, 2018.
[10] K. He, Y. Lu, and S. Sclaroff. Local descriptors optimized for average precision. Computer Vision
and Pattern Recognition, 2018.
[11] M. Dusmanu, I. Rocco, T. Pajdla, M. Pollefeys, J. Sivic, A. Torii, and T. Sattler. D2-net: A trainable cnn for joint description and detection of local features. Conference on Computer Vision and
Pattern Recognition, 2019.
[12] J. Revaud, C. De Souza, M. Humenberger, and P. Weinzaepfel. R2d2: Reliable and repeatable
detector and descriptor. Neural Information Processing Systems, 2019.
[13] O. Simeoni, Y. Avrithis, and O. Chum. Local features and visual words emerge in activations.
Conference on Computer Vision and Pattern Recognition, 2019.
[14] Z. Zhang, L. Wang, L. Zhou, and P. Koniusz. Learning spatial-context-aware global visual feature
representation for instance image retrieval. International Conference on Computer Vision, 2023.
[15] H. Wu, M. Wang, W. Zhou, Z. Lu, and H. Li. Asymmetric feature fusion for image retrieval. 2023.
[16] N. Vo, L. Jiang, C. Sun, K. Murphy, L.-J. Li, L. Fei-Fei, and J. Hays. Composing text and image for
image retrieval - an empirical odyssey. Computer Vision and Pattern Recognition, 2019.
[17] Z. Liu, C. Rodriguez-Opazo, D. Teney, and S. Gould. Image retrieval on real-life images with
pre-trained vision-and-language models. International Conference on Computer Vision, 2021.
[18] A. Baldrati, M. Bertini, T. Uricchio, and A. Del Bimbo. Conditioned and composed image retrieval
combining and partially fine-tuning clip-based features. Conference on Computer Vision and
Pattern Recognition, 2022.
[19] Y. Tian, S. Newsam, and K. Boakye. Fashion image retrieval with text feedback by additive attention
compositional learning. Conference on Applications of Computer Vision, 2023.