Recent advancements in large language models (LLMs) have greatly enhanced their ability to generate natural and contextually relevant text, enabling more human-like AI interactions. However, generating and understanding interactive human-like motion, where multiple individuals engage in coordinated movements, remains challenging due to the complexity of modeling these coordinated interactions. Furthermore, a unified and versatile model is required to handle diverse interactive scenarios, such as chat systems that dynamically adapt to user instructions and assigned roles. To tackle these problems, we introduce $\mathbf{M o L a M}$, the Interactive Motion-LAnguage Model, which integrates both language and motion modalities to effectively understand, generate, and control interactive motions in multi-turn conversational contexts. Unlike previous studies primarily focusing on uni-directional tasks (e.g., text-to-motion or motion-to-text), MoLaM employs a unified architecture capable of simultaneously understanding and generating both motion and text modalities. Given the lack of an appropriate dataset to address this challenge, we introduce Inter-MT2, a large-scale instructiontuning dataset containing 82.7K multi-turn interactive motion instructions, spanning 153 K interactive motion samples. Inter-MT2 covers diverse instructional scenarios including editing, question answering, and story generation, with interactive motions leveraging off-the-shelf large language models and motion diffusion models. We extensively evaluate the versatility of $\mathbf{M o L} \boldsymbol{a} \mathbf{M}$ across multiple interactive motion-related tasks: motion-to-text, text-to-motion, reaction generation, motion editing, and reasoning about motion sequences. Remarkably, $\mathbf{M o L} \boldsymbol{a} \mathbf{M}$ is the first model capable of effectively addressing all these tasks with a single unified framework, achieving competitive performance compared to task-specific methods.