Learning Generalizable Human Motion Generator with Reinforcement Learning

Text-driven human motion generation, as one of the vital tasks in\ncomputer-aided content creation, has recently attracted increasing attention.\nWhile pioneering research has largely focused on improving numerical\nperformance metrics on given datasets, practical applications reveal a common\nchallenge: existing methods often overfit specific motion expressions in the\ntraining data, hindering their ability to generalize to novel descriptions like\nunseen combinations of motions. This limitation restricts their broader\napplicability. We argue that the aforementioned problem primarily arises from\nthe scarcity of available motion-text pairs, given the many-to-many nature of\ntext-driven motion generation. To tackle this problem, we formulate\ntext-to-motion generation as a Markov decision process and present\n\\textbf{InstructMotion}, which incorporate the trail and error paradigm in\nreinforcement learning for generalizable human motion generation. Leveraging\ncontrastive pre-trained text and motion encoders, we delve into optimizing\nreward design to enable InstructMotion to operate effectively on both paired\ndata, enhancing global semantic level text-motion alignment, and synthetic\ntext-only data, facilitating better generalization to novel prompts without the\nneed for ground-truth motion supervision. Extensive experiments on prevalent\nbenchmarks and also our synthesized unpaired dataset demonstrate that the\nproposed InstructMotion achieves outstanding performance both quantitatively\nand qualitatively.\n

Paper

References (53)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC