Weaknesses
My concerns and suggestions of this manuscript are listed as blew:
1. The motivation of this work sounds unclear and looks a little incremental. As claimed in the abstract section 'the high computational complexity of these techniques limits the application in reality'. This statement is not clear. Does computational complexity mean more model parameters? or training time? or inference time? or more GPU memory cost? The authors didn't have a clear statement.
2. Moreover, in the abstract, 'which stem primarily from the diverse movement dynamics of various body parts.' are actually the general difficulties of the co-speech gesture generation task, not the specific one of how to employ the SSMs to this task.
3. The introduction section is not clear. The authors pay much more attention to the related works. Not the motivation and the high-level technical contributions of their work. This led me to feel that the entire work lacked technological innovation after reading the Introduction.
4. As for the methods, it is just directly imply the SSMs to this work. I cannot see any design on how to effectively solve the problem of 'computational complexity'. Actually, the technical contribution is poor, and the overall pipeline is very similar to the previous work EMAGE[1].
5. Could the authors explain why they only experimented on the BEATX dataset? As far as I know, TED[2] and TED-expressive[3] are also two commonly used co-speech gesture datasets.
6. The author did not conduct experiments to demonstrate how to effectively reduce the computational complexity by using SSMs, which directly led to the unclear motivation of this work.
[1] Liu, H., Zhu, Z., Becherini, G., Peng, Y., Su, M., Zhou, Y., ... & Black, M. J. (2023). Emage: Towards unified holistic co-speech gesture generation via masked audio gesture modeling. arXiv preprint arXiv:2401.00374.
[2] Yoon, Y., Cha, B., Lee, J. H., Jang, M., Lee, J., Kim, J., & Lee, G. (2020). Speech gesture generation from the trimodal context of text, audio, and speaker identity. ACM Transactions on Graphics (TOG), 39(6), 1-16.
[3] Liu, X., Wu, Q., Zhou, H., Xu, Y., Qian, R., Lin, X., ... & Zhou, B. (2022). Learning hierarchical cross-modal association for co-speech gesture generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10462-10472).