Video Prediction with Temporal-Spatial Attention Mechanism and Deep Perceptual Similarity Branch

Video prediction is a challenging but worth exploring task in computer vision. Different from image analysis, the challenge of video analysis derives from more complicated dependencies in time as well as in space. In this paper, we propose a Temporal-Spatial Attention Mechanism (TSAM) to capture not only spatial appearance dependencies but also temporal dynamic dependencies in video sequence. The TSAM is transplantable for existing networks and allows the long-range dependency modeling for various video analysis tasks (we take video prediction for specific experiments in this paper). Besides, we propose an additional Deep Perceptual Similarity Branch (DPSB) to encourage a better approximation to the ground-truth in high-level feature space, laying the foundation for frame generation. Extensive experiments on KTH, Penn Action and UCF-101 datasets demonstrate that our model performs quite competitively across diverse natural visual scenes, even in long-term video prediction.

Paper

Full text

PDF

Video Prediction with Temporal-Spatial Attention Mechanism and Deep Perceptual Similarity Branch

Semantic Scholar · Computer Science · 2019

Abstract

Video prediction is a challenging but worth exploring task in computer vision. Different from image analysis, the challenge of video analysis derives from more complicated dependencies in time as well as in space. In this paper, we propose a Temporal-Spatial Attention Mechanism (TSAM) to capture not only spatial appearance dependencies but also temporal dynamic dependencies in video sequence. The TSAM is transplantable for existing networks and allows the long-range dependency modeling for various video analysis tasks (we take video prediction for specific experiments in this paper). Besides, we propose an additional Deep Perceptual Similarity Branch (DPSB) to encourage a better approximation to the ground-truth in high-level feature space, laying the foundation for frame generation. Extensive experiments on KTH, Penn Action and UCF-101 datasets demonstrate that our model performs quite competitively across diverse natural visual scenes, even in long-term video prediction.

Similar papers

© 2026 NYSGPT2525 LLC