Video action detection requires dense spatio-temporal anno-tations which are challenging as well as expensive to obtain. However, real-world videos often have varying level of dif-ficulty and may not require equal level of annotations. In this paper we analyze the types of annotation appropriate for each sample and how it affects spatio-temporal video action detection. We focus on two different aspects affecting video action detection; 1) how to obtain varying level of annotations for videos, and 2) how to learn video action detection with different types of annotations. We study several annotation types including i) video level tags, ii) points iii) scribbles, iv) bounding box, and v) pixel level masks. First, we propose a simple active learning strategy which estimates appropriate types of annotations required for each video sample. Next, we propose a novel learning based spatio-temporal 3D-superpixel approach which gen-erates pseudo-labels from different types of annotations and enables learning of video action detection from such anno-tations. We validate our approach on two different datasets, UCFIOl-24 and IHMDB-21, for video action detection, sig-nificantly reducing the annotation cost without significant drop in performance.
Paper
References (82)
Scroll for more · 38 remaining