FasterVideo: Efficient Online Joint Object Detection And Tracking

. Object detection and tracking in videos represent essential and computationally demanding building blocks for current and future visual perception systems. In order to reduce the efficiency gap between available methods and computational requirements of real-world applications, we propose to re-think one of the most successful methods for image object detection, Faster R-CNN, and extend it to the video domain. Specifically, we extend the detection framework to learn instance-level embeddings which prove beneficial for data association and re-identification purposes. Focusing on the computational aspects of detection and tracking, our proposed method reaches a very high computational efficiency necessary for relevant applications, while still managing to compete with recent and state-of-the-art methods as shown in the experiments we conduct on standard object tracking benchmarks 3 . com-pares to other vision based methods published in the literature. For all methods not incorporating detection time in their performance evaluation we added the cost of Faster R-CNN. Results show that our proposed method is able to compete with other well performing methods, while, at the same time, achieving near real-time inference for both detection and tracking, highlighting the advantage of addressing the tasks jointly. Our method, additionally, is fully online, and requires no additional labels.

Paper

References (37)

Scroll for more · 25 remaining

Similar papers

© 2026 NYSGPT2525 LLC