End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding

Natural language spatial video grounding aims to detect the relevant objects\nin video frames with descriptive sentences as the query. In spite of the great\nadvances, most existing methods rely on dense video frame annotations, which\nrequire a tremendous amount of human effort. To achieve effective grounding\nunder a limited annotation budget, we investigate one-shot video grounding, and\nlearn to ground natural language in all video frames with solely one frame\nlabeled, in an end-to-end manner. One major challenge of end-to-end one-shot\nvideo grounding is the existence of videos frames that are either irrelevant to\nthe language query or the labeled frames. Another challenge relates to the\nlimited supervision, which might result in ineffective representation learning.\nTo address these challenges, we designed an end-to-end model via Information\nTree for One-Shot video grounding (IT-OS). Its key module, the information\ntree, can eliminate the interference of irrelevant frames based on branch\nsearch and branch cropping techniques. In addition, several self-supervised\ntasks are proposed based on the information tree to improve the representation\nlearning under insufficient labeling. Experiments on the benchmark dataset\ndemonstrate the effectiveness of our model.\n

Paper

References (64)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC