Interpreting video features: a comparison of 3D convolutional networks and convolutional LSTM networks

A number of techniques for interpretability have been presented for deep\nlearning in computer vision, typically with the goal of understanding what the\nnetworks have based their classification on. However, interpretability for deep\nvideo architectures is still in its infancy and we do not yet have a clear\nconcept of how to decode spatiotemporal features. In this paper, we present a\nstudy comparing how 3D convolutional networks and convolutional LSTM networks\nlearn features across temporally dependent frames. This is the first comparison\nof two video models that both convolve to learn spatial features but have\nprincipally different methods of modeling time. Additionally, we extend the\nconcept of meaningful perturbation introduced by \\cite{MeaningFulPert} to the\ntemporal dimension, to identify the temporal part of a sequence most meaningful\nto the network for a classification decision. Our findings indicate that the 3D\nconvolutional model concentrates on shorter events in the input sequence, and\nplaces its spatial focus on fewer, contiguous areas.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC