Joint Visual-Temporal Embedding for Unsupervised Learning of Actions in Untrimmed Sequences

Understanding the structure of complex activities in untrimmed videos is a\nchallenging task in the area of action recognition. One problem here is that\nthis task usually requires a large amount of hand-annotated minute- or even\nhour-long video data, but annotating such data is very time consuming and can\nnot easily be automated or scaled. To address this problem, this paper proposes\nan approach for the unsupervised learning of actions in untrimmed video\nsequences based on a joint visual-temporal embedding space. To this end, we\ncombine a visual embedding based on a predictive U-Net architecture with a\ntemporal continuous function. The resulting representation space allows\ndetecting relevant action clusters based on their visual as well as their\ntemporal appearance. The proposed method is evaluated on three standard\nbenchmark datasets, Breakfast Actions, INRIA YouTube Instructional Videos, and\n50 Salads. We show that the proposed approach is able to provide a meaningful\nvisual and temporal embedding out of the visual cues present in contiguous\nvideo frames and is suitable for the task of unsupervised temporal segmentation\nof actions.\n

Paper

References (41)

Scroll for more · 29 remaining

Similar papers

© 2026 NYSGPT2525 LLC