FashionSearchNet-v2: Learning Attribute Representations with Localization for Image Retrieval with Attribute Manipulation

The focus of this paper is on the problem of image retrieval with attribute\nmanipulation. Our proposed work is able to manipulate the desired attributes of\nthe query image while maintaining its other attributes. For example, the collar\nattribute of the query image can be changed from round to v-neck to retrieve\nsimilar images from a large dataset. A key challenge in e-commerce is that\nimages have multiple attributes where users would like to manipulate and it is\nimportant to estimate discriminative feature representations for each of these\nattributes. The proposed FashionSearchNet-v2 architecture is able to learn\nattribute specific representations by leveraging on its weakly-supervised\nlocalization module, which ignores the unrelated features of attributes in the\nfeature space, thus improving the similarity learning. The network is jointly\ntrained with the combination of attribute classification and triplet ranking\nloss to estimate local representations. These local representations are then\nmerged into a single global representation based on the instructed attribute\nmanipulation where desired images can be retrieved with a distance metric. The\nproposed method also provides explainability for its retrieval process to help\nprovide additional information on the attention of the network. Experiments\nperformed on several datasets that are rich in terms of the number of\nattributes show that FashionSearchNet-v2 outperforms the other state-of-the-art\nattribute manipulation techniques. Different than our earlier work\n(FashionSearchNet), we propose several improvements in the learning procedure\nand show that the proposed FashionSearchNet-v2 can be generalized to different\ndomains other than fashion.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC