# Response to reviewer `#bzXT` (1/2)
We appreciate the reviewer's inquiry regarding the comparison with SMPL-based 4D HOI generation methods. We acknowledge the valuable contributions made by existing methods in advancing human motion generation.
Our approach builds upon the strengths and robustness exhibited by these methods, aiming to enhance the realism and usability of 4D HOI generation. With a better 4D HOI method, our approach will also show improved performance. We hope for this paper to not only showcase the potential of 4D HOI but also to potentially advance and elevate our field.
> a substantial body of related work in HOI generation, particularly in 4D HOI generation methods. These methods have contributed significantly to the field by generating SMPL and SMPL-X poses for HOI scenarios.
> How does AvatarGo compare to SMPL-based 4D HOI generation methods?
We concur with the reviewer's observation that Programmable Motion Generation, Chain-of-Contacts, CG-HOI, and HSI have significantly advanced the field. However, these advancements primarily focus on generating motion sequences for SMPL or SMPL-X models, which lack intricate clothing details. Additionally, these methods rely on training datasets, limiting their adaptability in practical scenarios. While InterDreamer recently introduced a zero-shot framework for generating appearance within 4D HOI, their models still closely resemble the original SMPL design, resulting in minimal clothing details in their outputs. In contrast, our approach offers a novel avenue for zero-shot text-guided 4D HOI generation, producing both realistic appearance and geometry enhancements.
> While the paper focuses on generating realistic appearances using Gaussian splatting, which differs from the approaches of these methods, it is essential for the authors to analyze and position their work in the context of existing studies in 4D HOI generation.
Similar to the above discussion, our approach presents a novel opportunity for zero-shot text-guided 4D HOI generation, emphasizing the generation of realistic appearance and geometry. We will include more discussion with these 4D HOI generation methods and highlight the importance of the SMPL-based 4D HOI methods within this research domain.
> Another weakness of the proposed method, as mentioned by the authors in the limitation section, is that it assumes a consistent geometric relationship between the interacting body part and the object.
We agree with the reviewer on this point. However, We believe our approach to handling rigid-body interactions and continuous contact in 4D human animation represents a significant advancement. Meanwhile, this scenario, where the generated avatar interacts with rigid objects, is common in real-world applications, yet existing methods struggle with it. In other words, our proposed method outperforms all existing related approaches. Moreover, similar applications in robotics, such as embodied AI, primarily focus on resolving interactions between rigid structures. Similarly, SAGA[1] and other works focus only on a small range of settings, such as approaching and grasping an object. As an early study on generating 4D HOI scenes, we believe our method is valuable to the research community and will inspire future studies addressing more challenging cases. We will also aim to solve this problem in the future works.
> Aside from the apparent appearance advantage, does AvatarGo achieve a similar interaction quality to the SMPL-based methods?
Our methodology is centered around continuous interactions between humans and objects, making direct comparisons with SAGA[1] and similar works that concentrate on object grasping somewhat unfair.
When comparing with SMPL-based approaches operating within a similar context, we would like to direct the reviewer to the webpage of InterDreamer[2] (their code is not publicly released for comparison). It is noticeable that this method encounters challenges in maintaining realistic human-object interactions, often leading to insufficient contact or significant penetration. These observations demonstrate that our method achieves comparable or superior performance when compared to existing SMPL-based techniques.