Skeleton-based action recognition has recently gained profound attention due to its robust nature towards scene variation and illumination change. In this deep learning era, the increasing popularity of Convolutional Neural Network (CNN) in almost every computer vision task has made the researcher inspect its significance over graph-based skeleton data. Graph Convolutional Network (GCN) has made convolution operation available for joints (nodes) and bones (edges) of the human body during any action. An automatic Spatio-Temporal GCN (ST-GCN) is introduced instead of conventional handcrafted and traversal rule for human joints and bones modeling to move beyond assertion of confined expressive power and generalization capacity. This work investigates the cross-action transductive transfer learning technique on the graph using STGCTN architecture that combines ST-GCN and transfer learning together. We use the spatial configuration partitioning technique to model the graph data before sending it to the ST-GCTN model for spatial and temporal data exploitation. This parameter-based transfer learning technique is not only responsible for scaling down the training data volume but also enhances the expressive power to a greater extent. The proposed transfer learning-based graph model (ST-GCTN) outperforms many state-of-the-art STGCN methods by a significant margin by extensively transferring knowledge between two sets of actions inside a large-scale NTU RGB+D 60 dataset. The superior performance on the NW-UCLA dataset confirms our model's generalizability.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex