Multimodal feature fusion for CNN-based gait recognition: an empirical comparison

People identification in video based on the way they walk (i.e. gait) is a\nrelevant task in computer vision using a non-invasive approach. Standard and\ncurrent approaches typically derive gait signatures from sequences of binary\nenergy maps of subjects extracted from images, but this process introduces a\nlarge amount of non-stationary noise, thus, conditioning their efficacy. In\ncontrast, in this paper we focus on the raw pixels, or simple functions derived\nfrom them, letting advanced learning techniques to extract relevant features.\nTherefore, we present a comparative study of different Convolutional Neural\nNetwork (CNN) architectures by using three different modalities (i.e. gray\npixels, optical flow channels and depth maps) on two widely-adopted and\nchallenging datasets: TUM-GAID and CASIA-B. In addition, we perform a\ncomparative study between different early and late fusion methods used to\ncombine the information obtained from each kind of modalities. Our experimental\nresults suggest that (i) the raw pixel values represent a competitive input\nmodality, compared to the traditional state-of-the-art silhouette-based\nfeatures (e.g. GEI), since equivalent or better results are obtained; (ii) the\nfusion of the raw pixel information with information from optical flow and\ndepth maps allows to obtain state-of-the-art results on the gait recognition\ntask with an image resolution several times smaller than the previously\nreported results; and, (iii) the selection and the design of the CNN\narchitecture are critical points that can make a difference between\nstate-of-the-art results or poor ones.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC