MVFuseNet: Improving End-to-End Object Detection and Motion Forecasting through Multi-View Fusion of LiDAR Data
In this work, we propose \\textit{MVFuseNet}, a novel end-to-end method for\njoint object detection and motion forecasting from a temporal sequence of LiDAR\ndata. Most existing methods operate in a single view by projecting data in\neither range view (RV) or bird's eye view (BEV). In contrast, we propose a\nmethod that effectively utilizes both RV and BEV for spatio-temporal feature\nlearning as part of a temporal fusion network as well as for multi-scale\nfeature learning in the backbone network. Further, we propose a novel\nsequential fusion approach that effectively utilizes multiple views in the\ntemporal fusion network. We show the benefits of our multi-view approach for\nthe tasks of detection and motion forecasting on two large-scale self-driving\ndata sets, achieving state-of-the-art results. Furthermore, we show that\nMVFusenet scales well to large operating ranges while maintaining real-time\nperformance.\n