In recent years, although significant progress has been made in text-driven motion generation methods, notable limitations remain in trajectory control accuracy and conditional fusion. Some approaches rely on post-processing strategies for trajectory adjustment, which can easily lead to semantic deviation; others attempt to integrate text and trajectory conditions but struggle with motion plausibility and consistency. To address these issues, this paper proposes a trajectory-conditioned diffusion framework termed TCMD (Trajectory-Conditioned Motion Diffusion), which achieves joint control of text and trajectory while ensuring both physical and semantic coherence in generated motions. The proposed method incorporates two key innovations: (1) a spatio-temporal transformer perception module that serves as a trajectory condition controller integrated into the diffusion model, and (2) a two-stage training paradigm that effectively decouples semantic understanding from precise control. Specifically, the trajectory semantic understanding module is pre-trained first and then embedded into the diffusion generation process via a ControlNet structure for injecting trajectory conditions. Experimental results on the HumanML3D dataset demonstrate that TCMD outperforms existing methods in several aspects, including trajectory accuracy, text-match quality, motion generation realism, and action diversity. These findings validate the effectiveness and advancement of the proposed approach for high-precision motion generation under multi-condition constraints.
Paper
Full text
TCMD: trajectory-conditioned motion diffusion model
Semantic Scholar · Computer Science · 2026
Abstract
In recent years, although significant progress has been made in text-driven motion generation methods, notable limitations remain in trajectory control accuracy and conditional fusion. Some approaches rely on post-processing strategies for trajectory adjustment, which can easily lead to semantic deviation; others attempt to integrate text and trajectory conditions but struggle with motion plausibility and consistency. To address these issues, this paper proposes a trajectory-conditioned diffusion framework termed TCMD (Trajectory-Conditioned Motion Diffusion), which achieves joint control of text and trajectory while ensuring both physical and semantic coherence in generated motions. The proposed method incorporates two key innovations: (1) a spatio-temporal transformer perception module that serves as a trajectory condition controller integrated into the diffusion model, and (2) a two-stage training paradigm that effectively decouples semantic understanding from precise control. Specifically, the trajectory semantic understanding module is pre-trained first and then embedded into the diffusion generation process via a ControlNet structure for injecting trajectory conditions. Experimental results on the HumanML3D dataset demonstrate that TCMD outperforms existing methods in several aspects, including trajectory accuracy, text-match quality, motion generation realism, and action diversity. These findings validate the effectiveness and advancement of the proposed approach for high-precision motion generation under multi-condition constraints.