Temporal action segmentation is a topic of increasing interest, however,\nannotating each frame in a video is cumbersome and costly. Weakly supervised\napproaches therefore aim at learning temporal action segmentation from videos\nthat are only weakly labeled. In this work, we assume that for each training\nvideo only the list of actions is given that occur in the video, but not when,\nhow often, and in which order they occur. In order to address this task, we\npropose an approach that can be trained end-to-end on such data. The approach\ndivides the video into smaller temporal regions and predicts for each region\nthe action label and its length. In addition, the network estimates the action\nlabels for each frame. By measuring how consistent the frame-wise predictions\nare with respect to the temporal regions and the annotated action labels, the\nnetwork learns to divide a video into class-consistent regions. We evaluate our\napproach on three datasets where the approach achieves state-of-the-art\nresults.\n
Paper
References (38)
Scroll for more · 26 remaining