Simple means Faster: Real-Time Human Motion Forecasting in Monocular First Person Videos on CPU
We present a simple, fast, and light-weight RNN based framework for\nforecasting future locations of humans in first person monocular videos. The\nprimary motivation for this work was to design a network which could accurately\npredict future trajectories at a very high rate on a CPU. Typical applications\nof such a system would be a social robot or a visual assistance system for all,\nas both cannot afford to have high compute power to avoid getting heavier, less\npower efficient, and costlier. In contrast to many previous methods which rely\non multiple type of cues such as camera ego-motion or 2D pose of the human, we\nshow that a carefully designed network model which relies solely on bounding\nboxes can not only perform better but also predicts trajectories at a very high\nrate while being quite low in size of approximately 17 MB. Specifically, we\ndemonstrate that having an auto-encoder in the encoding phase of the past\ninformation and a regularizing layer in the end boosts the accuracy of\npredictions with negligible overhead. We experiment with three first person\nvideo datasets: CityWalks, FPL and JAAD. Our simple method trained on CityWalks\nsurpasses the prediction accuracy of state-of-the-art method (STED) while being\n9.6x faster on a CPU (STED runs on a GPU). We also demonstrate that our model\ncan transfer zero-shot or after just 15% fine-tuning to other similar datasets\nand perform on par with the state-of-the-art methods on such datasets (FPL and\nDTP). To the best of our knowledge, we are the first to accurately forecast\ntrajectories at a very high prediction rate of 78 trajectories per second on\nCPU.\n
Paper
References (47)
Scroll for more · 35 remaining