Gabriella: An Online System for Real-Time Activity Detection in Untrimmed Security Videos

Activity detection in security videos is a difficult problem due to multiple\nfactors such as large field of view, presence of multiple activities, varying\nscales and viewpoints, and its untrimmed nature. The existing research in\nactivity detection is mainly focused on datasets, such as UCF-101, JHMDB,\nTHUMOS, and AVA, which partially address these issues. The requirement of\nprocessing the security videos in real-time makes this even more challenging.\nIn this work we propose Gabriella, a real-time online system to perform\nactivity detection on untrimmed security videos. The proposed method consists\nof three stages: tubelet extraction, activity classification, and online\ntubelet merging. For tubelet extraction, we propose a localization network\nwhich takes a video clip as input and spatio-temporally detects potential\nforeground regions at multiple scales to generate action tubelets. We propose a\nnovel Patch-Dice loss to handle large variations in actor size. Our online\nprocessing of videos at a clip level drastically reduces the computation time\nin detecting activities. The detected tubelets are assigned activity class\nscores by the classification network and merged together using our proposed\nTubelet-Merge Action-Split (TMAS) algorithm to form the final action\ndetections. The TMAS algorithm efficiently connects the tubelets in an online\nfashion to generate action detections which are robust against varying length\nactivities. We perform our experiments on the VIRAT and MEVA (Multiview\nExtended Video with Activities) datasets and demonstrate the effectiveness of\nthe proposed approach in terms of speed (~100 fps) and performance with\nstate-of-the-art results. The code and models will be made publicly available.\n

Paper

References (35)

Scroll for more · 23 remaining

Similar papers

© 2026 NYSGPT2525 LLC