This paper explores the deep learning models aiming at two tasks, which are classifying objects and recognizing human action from a video. The deep learning models are the convolutional neural networks and long short-term memory network. For the action recognition, the optical flow is employed as the feature representation of movement on each video. The video data simulates one person doing either taking, returning, or browsing items on a shelf. From the experiments, the model achieve accuracy of 56.41% of accuracy for the object classification task and 76.92% for the action recognition.
Paper
Full text
Object and Human Action Recognition From Video Using Deep Learning Models
Semantic Scholar · Computer Science · 2019
Abstract
This paper explores the deep learning models aiming at two tasks, which are classifying objects and recognizing human action from a video. The deep learning models are the convolutional neural networks and long short-term memory network. For the action recognition, the optical flow is employed as the feature representation of movement on each video. The video data simulates one person doing either taking, returning, or browsing items on a shelf. From the experiments, the model achieve accuracy of 56.41% of accuracy for the object classification task and 76.92% for the action recognition.