MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling

Motion generation from discrete quantization offers many advantages over continuous regression, but at the cost of inevitable approximation errors. Previous methods usually quantize the entire body pose into one code, which not only faces the difficulty in encoding all joints within one vector but also loses the spatial relationship between different joints. Differently, in this work we quantize each individual joint into one vector, which i) simplifies the quantization process as the complexity associated with a single joint is markedly lower than that of the entire pose; ii) maintains a spatial-temporal structure that preserves both the spatial relationships among joints and the temporal movement patterns; iii) yields a 2D token map, which enables the application of various 2D operations widely used in 2D images. Grounded in the 2D motion quantization, we build a spatial-temporal modeling framework, where 2D joint VQVAE, temporal-spatial 2D masking technique, and spatial-temporal 2D attention are proposed to take advantage of spatial-temporal signals among the 2D tokens. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, with a 26.6% decrease of FID on HumanML3D and a 29.9% decrease on KIT-ML. Project page: https://aigc3d.github.io/mogents.

Paper

Similar papers

Peer review

Reviewer 1rhT6/10 · confidence 4/52024-06-25

Summary

This paper present a new method for text to motion generation. In this method the human motion is represented as 2D tokens in a codebook. This allow the authors to apply 2D operation on 3D motions and use a 2D masking strategy. The architecture is composed of a VAE to learn the codebook and of a Transformer to learn the relation between CLIP embedding of the text input and the corresponding codebook tokens using spatial-temporal attention. With this method the authors outperforms state of the art approaches quantitatively and qualitatively. An ablation study shows the effect of each component.

Strengths

The paper is clear and detailed. The mixed use of codebook, masking ans spatial-temporal transformer is interesting. The method outperforms the state of the art quantitatively and qualitatively. The ablation shows well the effect of each component.

Weaknesses

The Figure 3 b is not very clear It seems that the CLIP embedding is concatenated to the flattened motion embedding and then positional encoding is applied while the text say that the CLIP embedding is added after positional encoding. I also don't understand why positional encoding is added twice, one on the flattened vector and one time on the matrix. An ablation to show that this concatenation of token and text is the better than for example cross attention would have been welcome. Another interesting ablation would have been to see the performance of the model when computing the temporal and spatial attention in parallel instead of sequentially. It is not clear whether P is added only during attention or also on the inputs like base transformer. There is also no explanation as to why add positional encoding after computing the attention matrix. On several metric the ground truth is beaten but the paper does not provide explanation fro this. The paper does not describe how FID features are extracted ? The classifier free guidance should be mentioned in the main paper not just in the appendix. It is an important component. Regarding the motion editing figure : why is only one hand being raised with the Temporal-Spatial Editing while the temporal editing results shows both hands being raised. Since the plural is used in the prompt this would indicate that temporal editing is better.

Questions

It should be mentioned somewhere that j^i_t contains the 3 dimension of joint i. A user study would have been nice. Metrics are difficult to use on these more complex actions. It might be better to mention clip in the overview instead of waiting fro section 3.5.

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The limitations are very briefly addressed.

Reviewer 1rhT2024-08-09

Rating after rebuttal

The authors clarified the few thing that weren't clear to me. The user study is appreciated. I keep my original weak accept rating. It seems that the values reported inside the user study graph are wrong (102,73%).

Authorsrebuttal2024-08-13

Thanks for the reviewer's feedback

Thanks for the reviewer's feedback. We are pleased to see that our response has provided clarification on some points. Sorry for the confusion. When we state (102, 73%), we mean that there are 102 votes, which represent 73% of the total votes. We will revise this for clarity.

Reviewer SBXK6/10 · confidence 4/52024-07-02

Summary

This paper proposes an approach for text-conditioned motion generation. A common practice in this area is to use a quantized representation of human motion obtained with a VQ-VAE. However, most prior works represent the full body by a single token, which makes accurate reconstruction complicated. In this work, the authors propose a new way of quantizing human motion: a single token is associated with a single joint. Then, the motion can be represented as a 2D grid of indices corresponding to spatial and temporal dimensions. Using this new representation, this paper proposes to generate motions using masked generative modeling. In summary, the contributions of this paper are: - A new quantization of the human motion, representing each joint in a 2D map of tokens. - A masking strategy allowing the leverage of spatiotemporal information preserved by the proposed quantization. - A masked generative modeling strategy to generate new motions conditioned on text input.

Strengths

I would say that the main strength of this paper is not the novelty: masked generative modeling was already used for human motion generation [17]. However, this paper brings new components that make a lot of sense and seem to greatly impact the results. The bottleneck of prior works (1 pose = 1 token) is well-identified, and the proposed quantization strategy addresses this problem effectively by associating a token to each joint. In addition to improving the reconstruction after quantization, this representation proves useful as it preserves the spatiotemporal structure of the motion. Carefully designed operations (2D token masking, spatial-temporal motion Transformer) benefit from the proposed representation despite its higher dimension. Another strength of this paper is the evaluation. The comparisons follow the standard procedures and seem totally fair to other methods. Providing confidence intervals by running experiments multiple times is an excellent practice. Even for the evaluation of the quantization in Table 2, I find it very good that the authors decreased the codebook size of the introduced model for fair comparison with other methods. In addition to the 2 datasets widely used for comparisons, the appendix provides results on numerous datasets, which is appreciated to evaluate the model's generalization capability. The ablations are also satisfying, as they allow to evaluate the impact of the main introduced components.

Weaknesses

The main weakness of this paper is that it is difficult to understand the quantization of human motion: - L76 "each joint is quantized to an individual code of a VQ book": From my understanding, with the residual quantization, each joint is quantized to a sequence of indices; the final code is the sum of codes corresponding to those indices and associated codebooks. - Equation 1: Given L159, it seems that one joint in the input is converted to one token. Equation 1 suggests the same (the input of the encoder would be of dimension 3). And then L272, "Both the encoder and decoder are constructed from 2 convolutional residual blocks with a downscale of 4," so I really do not understand at all. Is there a spatiotemporal reduction? - Equation 2: This does not correspond to residual quantization. Maybe it is meant to simplify the understanding, but I find it very confusing. Globally, it is very difficult to understand how the method works until we reach section 3.5.2. For instance, until then, I did not understand the notion of a 2D map since the residual quantization would have made the grid 3-dimensional. I also wondered how the masking could encompass the depth of the quantization. Other minor issues include: - L125: It would be better to mention methods that represent a single pose (or human) with multiple tokens [a,b]. - The presentation of Table 1 is not optimal. Giving the dataset in the table instead of the caption would be more clear (like in [17]). Also, why are there no bold results for diversity? It may look like this is because some other methods have better results. [a] Geng, Z., Wang, C., Wei, Y., Liu, Z., Li, H., & Hu, H. (2023). Human pose as compositional tokens. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 660-671). [b] Fiche, G., Leglaive, S., Alameda-Pineda, X., Agudo, A., & Moreno-Noguer, F. (2023). VQ-HPS: Human Pose and Shape Estimation in a Vector-Quantized Latent Space. arXiv preprint arXiv:2312.08291.

Questions

- How does the quantization work? I would like to understand if there is a spatial-temporal reduction, as the information in the paper seems contradictory. - From my understanding in Section 3.3, there is no information about which joint is processed in the encoder (so the encoder processes in the same manner every joint). Can the VQ-VAE be considered a quantization of the 3D space as a learned grid? - From Figure 2 and other explanations, the pose seems flattened to correspond to a column in the 2D map. How are the joints organized so that the spatial structure of the pose is preserved? Is there a topological ordering?

Rating

6

Confidence

4

Soundness

4

Presentation

2

Contribution

3

Limitations

The section about the limitations is very short. I think that it could be improved by proposing more solutions to the quantization problem and other research perspectives. Otherwise, it sounds like the problem of motion generation is now completely solved. The authors say that this work has no societal impact at all. I would agree that the impact on society is very limited.

Authorsrebuttal2024-08-12

Thanks for the reviewer's feedback. We are pleased to see that our response has addressed most of the concerns. Regarding the spatial structure, the flattened order follows the HumanML3D dataset and is as follows: 'pelvis', 'right_hip', 'left_hip', 'spine1', 'right_knee', 'left_knee', 'spine2', 'right_ankle', 'left_ankle', 'spine3', 'right_foot', 'left_foot', 'neck', 'right_collar', 'left_collar', 'head', 'right_shoulder', 'left_shoulder', 'right_elbow', 'left_elbow', 'right_wrist', 'left_wrist'. These joints are arranged in order of proximity to the pelvis joint, from nearest to furthest. Sure, we will move the limitation to the annexes and discuss more.

Reviewer g7ot5/10 · confidence 5/52024-07-13

Summary

This paper proposes a novel approach to human motion generation by quantizing each joint into individual vectors, rather than encoding the entire body pose into one code. The key contributions are: (1) It quantizes each joint separately to preserve spatial relationships and simplify the encoding process. Then the motion sequence are organized into a 2D token map, akin to 2D images, allowing the use of 2D operations like convolution, positional encoding, and attention mechanisms. (2) It introduces a spatial-temporal 2D joint VQVAE to encode motion sequences into discrete codes and employs a temporal-spatial 2D masking strategy and a spatial-temporal 2D transformer to predict masked tokens.

Strengths

1. The paper introduces a novel joint-level quantization approach, addressing the complexity and spatial information loss issues seen in whole-body pose quantization.By organizing motion sequences into a 2D token map, the method takes advantage of powerful 2D image processing techniques, enhancing feature extraction and motion generation. 2. The integration of 2D joint VQVAE, temporal-spatial 2D masking, and spatial-temporal 2D attention forms a robust framework that effectively captures spatial-temporal dynamics in human motion. 3. Extensive experiments demonstrate the method's efficacy, outperforming previous state-of-the-art methods on key datasets. 4. The paper is well-written and easy to understand. The supp. mat. video provides comparisons with MoMask.

Weaknesses

1. Although the overall idea of joint-level quantization is interesting, I still have the concern of computational overhead. While motion representation is typically lightweight, the use of 2D code maps and spatial-temporal attention can introduce significant computational overhead, similar to image data processing. It would be beneficial to compare the inference speed of mainstream methods (e.g., MoMask, T2M-GPT, MLD, etc) to show that the state-of-the-art performance is achieved with comparable computational costs. 2. The experiments are conducted on relatively small datasets (HumanML3D and KIT-ML). To better validate the effectiveness of the proposed method, experiments on larger-scale datasets, such as Motion-X, would be advantageous.

Questions

Please see the weaknesses. Overall, the idea is interesting. My major concern is the extra computational cost of this method, which could be much larger than previous methods with body-level VQ-VAE yet it is not investigated in the paper.

Rating

5

Confidence

5

Soundness

3

Presentation

3

Contribution

3

Limitations

This paper has discusses the limitation of this paper: the approximation error in VQ-VAE and a larger dataset for training more accurate VQ-VAE. I think there is no potential negative societal impact in this paper.

Reviewer SBXK2024-08-09

Thanks to the authors for an insightful rebuttal addressing most of my concerns. I still lean towards accepting this paper. I have a doubt about the **spatial structure**. I understand that the joints' ordering is the same at training and inference and that there is a positional encoding. My question was more about the order of the joints once flattened: does it follow an order that preserves the structure of the skeleton (for instance, left shoulder -> left elbow -> left wrist, ...), or is the spatial information exclusively in the positional encoding? For the lack of space in the caption of Table 1 and the limitations, I would suggest moving the limitations to the annexes.

Reviewer g7ot2024-08-12

After I carefully read other reviews and the author rebuttal, I think this paper proposes an effective and efficient method for motion generation. My initial concerns about the efficiency and the results on Motion-X has also been resolved in the author rebuttal. I will keep my original rating and leaning to accept this paper.

Authorsrebuttal2024-08-14

Thanks for the reviewer's feedback

Thanks for the reviewer's feedback. We are pleased to see that our response has addressed the concerns. If there are no further concerns, please also consider raising the rating. Many thanks!

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC