Discovering Dynamic Salient Regions for Spatio-Temporal Graph Neural Networks

Graph Neural Networks are perfectly suited to capture latent interactions\nbetween various entities in the spatio-temporal domain (e.g. videos). However,\nwhen an explicit structure is not available, it is not obvious what atomic\nelements should be represented as nodes. Current works generally use\npre-trained object detectors or fixed, predefined regions to extract graph\nnodes. Improving upon this, our proposed model learns nodes that dynamically\nattach to well-delimited salient regions, which are relevant for a higher-level\ntask, without using any object-level supervision. Constructing these localized,\nadaptive nodes gives our model inductive bias towards object-centric\nrepresentations and we show that it discovers regions that are well correlated\nwith objects in the video. In extensive ablation studies and experiments on two\nchallenging datasets, we show superior performance to previous graph neural\nnetworks models for video classification.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC