Modeling Multi-Label Action Dependencies for Temporal Action Localization

Real-world videos contain many complex actions with inherent relationships\nbetween action classes. In this work, we propose an attention-based\narchitecture that models these action relationships for the task of temporal\naction localization in untrimmed videos. As opposed to previous works that\nleverage video-level co-occurrence of actions, we distinguish the relationships\nbetween actions that occur at the same time-step and actions that occur at\ndifferent time-steps (i.e. those which precede or follow each other). We define\nthese distinct relationships as action dependencies. We propose to improve\naction localization performance by modeling these action dependencies in a\nnovel attention-based Multi-Label Action Dependency (MLAD)layer. The MLAD layer\nconsists of two branches: a Co-occurrence Dependency Branch and a Temporal\nDependency Branch to model co-occurrence action dependencies and temporal\naction dependencies, respectively. We observe that existing metrics used for\nmulti-label classification do not explicitly measure how well action\ndependencies are modeled, therefore, we propose novel metrics that consider\nboth co-occurrence and temporal dependencies between action classes. Through\nempirical evaluation and extensive analysis, we show improved performance over\nstate-of-the-art methods on multi-label action localization\nbenchmarks(MultiTHUMOS and Charades) in terms of f-mAP and our proposed metric.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC