In this paper, we consider the problem of human action recognition in realistic scenes. Mainstream methods adopt local spatio-temporal patterns matching strategy that detect STIPs (Spatio-Temporal Interest Points) and compute features from raw video frames, and then classify the features into some predefined action categories. However, these researches have made certain assumptions seldom hold in the real-world environment such as small scale, without viewpoint changes and background motion. To obtain more effective detection and description for human action videos of complex scenes, we improve the performance of STIPs detection in realistic scenes by filtering out noises resulted from low resolution or cluttered background. In addition, we propose a new feature description method that uses 3D local self-similarities to compensate for motion estimation error due to camera motion. BoW (Bag-of-Words) model is applied to learn a codebook in training stage, and a nonlinear SVM classifier is used for action recognition finally. Experimental results show that our approach achieve superior performance on more challenge dataset (Youtube) in comparison to baseline method.
Paper
Full text
Action recognition in realistic scenes via local spatio-temporal representation
Semantic Scholar · Computer Science · 2014
Abstract
In this paper, we consider the problem of human action recognition in realistic scenes. Mainstream methods adopt local spatio-temporal patterns matching strategy that detect STIPs (Spatio-Temporal Interest Points) and compute features from raw video frames, and then classify the features into some predefined action categories. However, these researches have made certain assumptions seldom hold in the real-world environment such as small scale, without viewpoint changes and background motion. To obtain more effective detection and description for human action videos of complex scenes, we improve the performance of STIPs detection in realistic scenes by filtering out noises resulted from low resolution or cluttered background. In addition, we propose a new feature description method that uses 3D local self-similarities to compensate for motion estimation error due to camera motion. BoW (Bag-of-Words) model is applied to learn a codebook in training stage, and a nonlinear SVM classifier is used for action recognition finally. Experimental results show that our approach achieve superior performance on more challenge dataset (Youtube) in comparison to baseline method.
References (22)
Scroll for more · 10 remaining