We present a method for weakly-supervised action localization based on graph\nconvolutions. In order to find and classify video time segments that correspond\nto relevant action classes, a system must be able to both identify\ndiscriminative time segments in each video, and identify the full extent of\neach action. Achieving this with weak video level labels requires the system to\nuse similarity and dissimilarity between moments across videos in the training\ndata to understand both how an action appears, as well as the sub-actions that\ncomprise the action's full extent. However, current methods do not make\nexplicit use of similarity between video moments to inform the localization and\nclassification predictions. We present a novel method that uses graph\nconvolutions to explicitly model similarity between video moments. Our method\nutilizes similarity graphs that encode appearance and motion, and pushes the\nstate of the art on THUMOS '14, ActivityNet 1.2, and Charades for weakly\nsupervised action localization.\n
Paper
References (45)
Scroll for more · 33 remaining