StarNet: Joint Action-Space Prediction with Star Graphs and Implicit Global Frame Self-Attention
In this work, we present a novel multi-modal multi-agent trajectory\nprediction architecture, focusing on map and interaction modeling using graph\nrepresentation. For the purposes of map modeling, we capture rich topological\nstructure into vector-based star graphs, which enable an agent to directly\nattend to relevant regions along polylines that are used to represent the map.\nWe denote this architecture StarNet, and integrate it in a single-agent\nprediction setting. As the main result, we extend this architecture to joint\nscene-level prediction, which produces multiple agents' predictions\nsimultaneously. The key idea in joint-StarNet is integrating the awareness of\none agent in its own reference frame with how it is perceived from the points\nof view of other agents. We achieve this via masked self-attention. Both\nproposed architectures are built on top of the action-space prediction\nframework introduced in our previous work, which ensures kinematically feasible\ntrajectory predictions. We evaluate the methods on the interaction-rich inD and\nINTERACTION datasets, with both StarNet and joint-StarNet achieving\nimprovements over state of the art.\n