SL-DML: Signal Level Deep Metric Learning for Multimodal One-Shot Action Recognition

Recognizing an activity with a single reference sample using metric learning\napproaches is a promising research field. The majority of few-shot methods\nfocus on object recognition or face-identification. We propose a metric\nlearning approach to reduce the action recognition problem to a nearest\nneighbor search in embedding space. We encode signals into images and extract\nfeatures using a deep residual CNN. Using triplet loss, we learn a feature\nembedding. The resulting encoder transforms features into an embedding space in\nwhich closer distances encode similar actions while higher distances encode\ndifferent actions. Our approach is based on a signal level formulation and\nremains flexible across a variety of modalities. It further outperforms the\nbaseline on the large scale NTU RGB+D 120 dataset for the One-Shot action\nrecognition protocol by 5.6%. With just 60% of the training data, our approach\nstill outperforms the baseline approach by 3.7%. With 40% of the training data,\nour approach performs comparably well to the second follow up. Further, we show\nthat our approach generalizes well in experiments on the UTD-MHAD dataset for\ninertial, skeleton and fused data and the Simitate dataset for motion capturing\ndata. Furthermore, our inter-joint and inter-sensor experiments suggest good\ncapabilities on previously unseen setups.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC