Weaknesses
- Poor writing. I list some cases as follows. **In the rebuttal session, I suggest authors revise all issues and list them on the openreview platform. I will check them one by one.**
- In this submission, there are full of incorrect citation format. For example, in the first sentence of the introduction, Zhang et al. (2023a) should be (Zhang et al. 2023a). The authors seem not clear about the difference between `\cite` and `\citep`.
- Orphan row. It is unprofessional, which seems trying to make up 9 pages. Like the final line of the introduction.
- Colloquial expressions are unprofessional. Case: “**By the way**, our training stage … …”.
- SemanticBoost or Semantic Boost? This is an inconsistent statement.
- “Motiongpt: Human motion as a foreign language” has been accepted by NeurIPS 2023. Please revise the citation. **Please list all similar issues and revise them.**
- Incorrect use of quotation marks. For example, in page 4, “forward” should be ``forward''.
- Incorrect use of comma marks. The comma before "which" is wrong.
- “… the diffusion-based generative method, such as MDM Tevet et al. (2022), MotionDiffuse Zhang et al. (2022) and faster MLD Chen et al. (2022) …”. What does “faster” mean? Faster in inference or training? It is confusing.
- Missing explanation of equations. In Eq. 1, what do $T_i$ and $p^{mid-ear}$ mean? This makes readers confused. Does $Z$ mean $Z+$?
- Missing discussion on some motion diffusion models. I list some related work.
[1]: Xu, Sirui, et al. "InterDiff: Generating 3D Human-Object Interactions with Physics-Informed Diffusion." *ICCV,* 2023.
[2]: Yuan, Ye, et al. "Physdiff: Physics-guided human motion diffusion model." *ICCV, 2023*.
[3]: Chen, Ling-Hao, et al. "HumanMAC: Masked Motion Completion for Human Motion Prediction." ICCV, 2023.
[4]: Barquero, German, Sergio Escalera, and Cristina Palmero. "Belfusion: Latent diffusion for behavior-driven human motion prediction." *ICCV,* 2023.
[5]: Shafir, Yonatan, et al. "Human motion diffusion as a generative prior." *arXiv preprint arXiv:2303.01418* (2023).
- **Unfair comparison. (I flag this submission as an ethical issue on fairness concerns and should be reviewed by an Ethics Reviewer.)**
- As discussed in the “motion representations” part, the authors take the 269-dim motion representation, which is different from baselines. Authors should report all baseline results with 269-dim motion representation.
- Why is the 269-dim motion feature better than the 263-dim feature? Please provide experiment ablation to verify it.
- Authors do not provide mean and standard values of results. Baselines report the mean and standard values with 3 times repeating.
- In section 2.2, the authors claim that “However, the improved text annotations may not effectively align with the ground-truth motion data”. Any experimental results to support it?
- In ablation, authors introduce TS, HOS, and LFS metrics. Any explanation? Any citation? Why do not compare it with the baselines on Tab. 1?
- **Missing results on KIT-ML.**
- Why does the joint number affect the experiments? If so, the method proposed in the paper will have significant limitations. Because it only applies to specific skeleton structures.
- If it affects the experiment, why not use motion retargeting methods?
- Missing discussion on limitation.
- **Missing ablation.**
- In Tab. 2, authors show the ablations on three components respectively. No results on the following settings: none of three, and with two components. The combination of ablation experiments should be 8 groups. Any missing ablation should explain why.
- No visualization result to support the ablation in Tab. 2. This makes it unclear how exactly any of these components benefit the generated results.
- Authors do not provide any codes and appendix to support the reproduction. This makes it hard to follow. Note that this is the reason to reject a paper, but it reduces the value of a work to the community.
- **Marginal contribution.** The semantic enhancement method only supports providing the direction of motion, which is only a small part of the motion. Besides, motions also include action counting, body-part motion, and something else. If the augmented text cues can resolve these, I suggest authors provide the following cases.
- “A person jumps three times.”
- “A person jumps four times.”
- “A person lifts his left leg then walks.”
- “A person lifts his right leg then walks.”
- The figure of DEFE is confusing. Please detail the network architecture.
- For network architecture, the names are complex, which makes it hard to read. Additionally, authors sometimes use their full names and sometimes use their abbreviated names. The statement here is hard to follow.
Questions
Question:
- How many words are in the status table? Authors should list the implementation details as detailed as possible.
- What will the results be if there is no text enhancement?
- **I am very confused about why the R-Precision is higher than GT. This is my main concern. Do authors use the augmented text for the test? If so, can authors report the R-Precision with original texts?**
- Does the feature extractor for calculating R-Precision use the checkpoint provided by the HumanML3D paper or the model is trained by the authors?
---
## Rating
My concerns come from the unfair comparison, poor writing, missing experiments, contribution, and metrics. I hope these concerns can be resolved. Now, I think this submission is below the borderline. After the rebuttal and discussion, I will revise my rating to a clear rating (reject or accept).