For autonomous agents to successfully operate in real world, the ability to\nanticipate future motions of surrounding entities in the scene can greatly\nenhance their safety levels since potentially dangerous situations could be\navoided in advance. While impressive results have been shown on predicting each\nagent's behavior independently, we argue that it is not valid to consider road\nentities individually since transitions of vehicle states are highly coupled.\nMoreover, as the predicted horizon becomes longer, modeling prediction\nuncertainties and multi-modal distributions over future sequences will turn\ninto a more challenging task. In this paper, we address this challenge by\npresenting a multi-modal probabilistic prediction approach. The proposed method\nis based on a generative model and is capable of jointly predicting sequential\nmotions of each pair of interacting agents. Most importantly, our model is\ninterpretable, which can explain the underneath logic as well as obtain more\nreliability to use in real applications. A complicate real-world roundabout\nscenario is utilized to implement and examine the proposed method.\n