ST-ABN: Visual Explanation Taking into Account Spatio-temporal Information for Video Recognition
It is difficult for people to interpret the decision-making in the inference\nprocess of deep neural networks. Visual explanation is one method for\ninterpreting the decision-making of deep learning. It analyzes the\ndecision-making of 2D CNNs by visualizing an attention map that highlights\ndiscriminative regions. Visual explanation for interpreting the decision-making\nprocess in video recognition is more difficult because it is necessary to\nconsider not only spatial but also temporal information, which is different\nfrom the case of still images. In this paper, we propose a visual explanation\nmethod called spatio-temporal attention branch network (ST-ABN) for video\nrecognition. It enables visual explanation for both spatial and temporal\ninformation. ST-ABN acquires the importance of spatial and temporal information\nduring network inference and applies it to recognition processing to improve\nrecognition performance and visual explainability. Experimental results with\nSomething-Something datasets V1 \\& V2 demonstrated that ST-ABN enables visual\nexplanation that takes into account spatial and temporal information\nsimultaneously and improves recognition performance.\n