Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

This paper presents a new task, the grounding of spatio-temporal identifying\ndescriptions in videos. Previous work suggests potential bias in existing\ndatasets and emphasizes the need for a new data creation schema to better model\nlinguistic structure. We introduce a new data collection scheme based on\ngrammatical constraints for surface realization to enable us to investigate the\nproblem of grounding spatio-temporal identifying descriptions in videos. We\nthen propose a two-stream modular attention network that learns and grounds\nspatio-temporal identifying descriptions based on appearance and motion. We\nshow that motion modules help to ground motion-related words and also help to\nlearn in appearance modules because modular neural networks resolve task\ninterference between modules. Finally, we propose a future challenge and a need\nfor a robust system arising from replacing ground truth visual annotations with\nautomatic video object detector and temporal event localization.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC