Toward the goal of automatic production for sports broadcasts, a paramount\ntask consists in understanding the high-level semantic information of the game\nin play. For instance, recognizing and localizing the main actions of the game\nwould allow producers to adapt and automatize the broadcast production,\nfocusing on the important details of the game and maximizing the spectator\nengagement. In this paper, we focus our analysis on action spotting in soccer\nbroadcast, which consists in temporally localizing the main actions in a soccer\ngame. To that end, we propose a novel feature pooling method based on NetVLAD,\ndubbed NetVLAD++, that embeds temporally-aware knowledge. Different from\nprevious pooling methods that consider the temporal context as a single set to\npool from, we split the context before and after an action occurs. We argue\nthat considering the contextual information around the action spot as a single\nentity leads to a sub-optimal learning for the pooling module. With NetVLAD++,\nwe disentangle the context from the past and future frames and learn specific\nvocabularies of semantics for each subsets, avoiding to blend and blur such\nvocabulary in time. Injecting such prior knowledge creates more informative\npooling modules and more discriminative pooled features, leading into a better\nunderstanding of the actions. We train and evaluate our methodology on the\nrecent large-scale dataset SoccerNet-v2, reaching 53.4% Average-mAP for action\nspotting, a +12.7% improvement w.r.t the current state-of-the-art.\n
Paper
References (44)
Scroll for more · 32 remaining