Fusion-GCN: Multimodal Action Recognition using Graph Convolutional Networks

In this paper, we present Fusion-GCN, an approach for multimodal action\nrecognition using Graph Convolutional Networks (GCNs). Action recognition\nmethods based around GCNs recently yielded state-of-the-art performance for\nskeleton-based action recognition. With Fusion-GCN, we propose to integrate\nvarious sensor data modalities into a graph that is trained using a GCN model\nfor multi-modal action recognition. Additional sensor measurements are\nincorporated into the graph representation, either on a channel dimension\n(introducing additional node attributes) or spatial dimension (introducing new\nnodes). Fusion-GCN was evaluated on two public available datasets, the\nUTD-MHAD- and MMACT datasets, and demonstrates flexible fusion of RGB\nsequences, inertial measurements and skeleton sequences. Our approach gets\ncomparable results on the UTD-MHAD dataset and improves the baseline on the\nlarge-scale MMACT dataset by a significant margin of up to 12.37% (F1-Measure)\nwith the fusion of skeleton estimates and accelerometer measurements.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC