Not 3D Re-ID: a Simple Single Stream 2D Convolution for Robust Video Re-identification

Video-based person re-identification has received increasing attention\nrecently, as it plays an important role within surveillance video analysis.\nVideo-based Re-ID is an expansion of earlier image-based re-identification\nmethods by learning features from a video via multiple image frames for each\nperson. Most contemporary video Re-ID methods utilise complex CNNbased network\narchitectures using 3D convolution or multibranch networks to extract\nspatial-temporal video features. By contrast, in this paper, we illustrate\nsuperior performance from a simple single stream 2D convolution network\nleveraging the ResNet50-IBN architecture to extract frame-level features\nfollowed by temporal attention for clip level features. These clip level\nfeatures can be generalised to extract video level features by averaging\nwithout any significant additional cost. Our approach uses best video Re-ID\npractice and transfer learning between datasets to outperform existing\nstate-of-the-art approaches on the MARS, PRID2011 and iLIDS-VID datasets with\n89:62%, 97:75%, 97:33% rank-1 accuracy respectively and with 84:61% mAP for\nMARS, without reliance on complex and memory intensive 3D convolutions or\nmulti-stream networks architectures as found in other contemporary work.\nConversely, our work shows that global features extracted by the 2D convolution\nnetwork are a sufficient representation for robust state of the art video\nRe-ID.\n

Paper

References (40)

Scroll for more · 28 remaining

Similar papers

© 2026 NYSGPT2525 LLC