The state-of-the art solutions for human activity understanding from a video\nstream formulate the task as a spatio-temporal problem which requires joint\nlocalization of all individuals in the scene and classification of their\nactions or group activity over time. Who is interacting with whom, e.g. not\neveryone in a queue is interacting with each other, is often not predicted.\nThere are scenarios where people are best to be split into sub-groups, which we\ncall social groups, and each social group may be engaged in a different social\nactivity. In this paper, we solve the problem of simultaneously grouping people\nby their social interactions, predicting their individual actions and the\nsocial activity of each social group, which we call the social task. Our main\ncontributions are: i) we propose an end-to-end trainable framework for the\nsocial task; ii) our proposed method also sets the state-of-the-art results on\ntwo widely adopted benchmarks for the traditional group activity recognition\ntask (assuming individuals of the scene form a single group and predicting a\nsingle group activity label for the scene); iii) we introduce new annotations\non an existing group activity dataset, re-purposing it for the social task.\n
Paper
References (71)
Scroll for more · 38 remaining