In this paper, we tackle the problem of detecting objects in 3D and\nforecasting their future motion in the context of self-driving. Towards this\ngoal, we design a novel approach that explicitly takes into account the\ninteractions between actors. To capture their spatial-temporal dependencies, we\npropose a recurrent neural network with a novel Transformer architecture, which\nwe call the Interaction Transformer. Importantly, our model can be trained\nend-to-end, and runs in real-time. We validate our approach on two challenging\nreal-world datasets: ATG4D and nuScenes. We show that our approach can\noutperform the state-of-the-art on both datasets. In particular, we\nsignificantly improve the social compliance between the estimated future\ntrajectories, resulting in far fewer collisions between the predicted actors.\n
Paper
References (48)
Scroll for more · 36 remaining