Joint Representation of Temporal Image Sequences and Object Motion for Video Object Detection
In this paper, we propose a new video object detector (VoD) method referred\nto as temporal feature aggregation and motion-aware VoD (TM-VoD), which\nproduces a joint representation of temporal image sequences and object motion.\nThe proposed TM-VoD aggregates visual feature maps extracted by convolutional\nneural networks applying the temporal attention gating and spatial feature\nalignment. This temporal feature aggregation is performed in two stages in a\nhierarchical fashion. In the first stage, the visual feature maps are fused at\na pixel level via gated attention model. In the second stage, the proposed\nmethod aggregates the features after aligning the object features using\ntemporal box offset calibration and weights them according to the cosine\nsimilarity measure. The proposed TM-VoD also finds the representation of the\nmotion of objects in two successive steps. The pixel-level motion features are\nfirst computed based on the incremental changes between the adjacent visual\nfeature maps. Then, box-level motion features are obtained from both the region\nof interest (RoI)-aligned pixel-level motion features and the sequential\nchanges of the box coordinates. Finally, all these features are concatenated to\nproduce a joint representation of the objects for VoD. The experiments\nconducted on the ImageNet VID dataset demonstrate that the proposed method\noutperforms existing VoD methods and achieves a performance comparable to that\nof state-of-the-art VoDs.\n
Paper
References (30)
Scroll for more · 18 remaining