Efficient Text-based Person Search via Single-stage Identity-guided Attribute Parsing and Alignment

Cross-modal text-based person search aims at retrieving target person in a large image gallery by natural language description. This task is quite challenging due to the complex environment of person image acquisition and the semantic gap between different modalities. The existing popular deep learning based models rely heavily on a large amount of labeled data to obtain good performance, which is labor-consuming and not always available in real applications. In order to achieve an effective alignment within and between modalities, additional semantic information or pre-trained network is often introduced to assist human body positioning, which will further shape the increase in training parameters and the decrease in efficiency. To address these problems, we propose a single-stage Identity-guided image-text Attribute Parsing and Alignment network (IAPA). IAPA realizes cross-modal alignment of human body parts in an unsupervised manner through image itself, resulting in great efficiency improvement while maintaining promising accuracy. This is also the first attempt to apply pixel-level supervision to cross-modal person retrieval task. The experiments on the CUHK-PEDES data-set validate the effectiveness and the efficiency of IAPA compared to the other state-of-the-art methods.

Paper

Full text

PDF

Efficient Text-based Person Search via Single-stage Identity-guided Attribute Parsing and Alignment

OpenAlex · Video Surveillance and Tracking Methods · 2022

Abstract

Cross-modal text-based person search aims at retrieving target person in a large image gallery by natural language description. This task is quite challenging due to the complex environment of person image acquisition and the semantic gap between different modalities. The existing popular deep learning based models rely heavily on a large amount of labeled data to obtain good performance, which is labor-consuming and not always available in real applications. In order to achieve an effective alignment within and between modalities, additional semantic information or pre-trained network is often introduced to assist human body positioning, which will further shape the increase in training parameters and the decrease in efficiency. To address these problems, we propose a single-stage Identity-guided image-text Attribute Parsing and Alignment network (IAPA). IAPA realizes cross-modal alignment of human body parts in an unsupervised manner through image itself, resulting in great efficiency improvement while maintaining promising accuracy. This is also the first attempt to apply pixel-level supervision to cross-modal person retrieval task. The experiments on the CUHK-PEDES data-set validate the effectiveness and the efficiency of IAPA compared to the other state-of-the-art methods.

Similar papers

© 2026 NYSGPT2525 LLC