The automatic characterization of pedestrians in surveillance footage is a\ntough challenge, particularly when the data is extremely diverse with cluttered\nbackgrounds, and subjects are captured from varying distances, under multiple\nposes, with partial occlusion. Having observed that the state-of-the-art\nperformance is still unsatisfactory, this paper provides a novel solution to\nthe problem, with two-fold contributions: 1) considering the strong semantic\ncorrelation between the different full-body attributes, we propose a multi-task\ndeep model that uses an element-wise multiplication layer to extract more\ncomprehensive feature representations. In practice, this layer serves as a\nfilter to remove irrelevant background features, and is particularly important\nto handle complex, cluttered data; and 2) we introduce a weighted-sum term to\nthe loss function that not only relativizes the contribution of each task (kind\nof attributed) but also is crucial for performance improvement in\nmultiple-attribute inference settings. Our experiments were performed on two\nwell-known datasets (RAP and PETA) and point for the superiority of the\nproposed method with respect to the state-of-the-art. The code is available at\nhttps://github.com/Ehsan-Yaghoubi/MAN-PAR-.\n
Paper
References (43)
Scroll for more · 31 remaining