We create a family of powerful video models which are able to: (i) learn\ninteractions between semantic object information and raw appearance and motion\nfeatures, and (ii) deploy attention in order to better learn the importance of\nfeatures at each convolutional block of the network. A new network component\nnamed peer-attention is introduced, which dynamically learns the attention\nweights using another block or input modality. Even without pre-training, our\nmodels outperform the previous work on standard public activity recognition\ndatasets with continuous videos, establishing new state-of-the-art. We also\nconfirm that our findings of having neural connections from the object modality\nand the use of peer-attention is generally applicable for different existing\narchitectures, improving their performances. We name our model explicitly as\nAssembleNet++. The code will be available at:\nhttps://sites.google.com/corp/view/assemblenet/\n
Paper
References (62)
Scroll for more · 38 remaining