Thanks to the reviewers for the time and effort dedicated to thoroughly reviewing our work.
**As a 'training-free' or 'plug-in' module, I cannot see the detailed generalization examples of the proposed method to other multi-view 3D generation and video generation methods. This cannot support the claim in the paper.**
Thanks for the comments. We first clarify the following facets:
(i) For the 'training-free' property, our method indeed only utilizes the off-the-shelf multi-view image generative model and the video generative model without additional training.
(ii) Although we did not claim the generalization ability implied by "plug-in" as our contribution in the paper for rigorousness, this property is evident in logic, as the proposed pipeline holds no bias on the model architecture. Therefore, we give an affirmative response to Q4.
Considering the above points, we use SV3D [4] as the multi-view 3D generator given its superior performance than other MVDiffusions.
It is sufficient to demonstrate the feasibility of our proposed approach.
However, to solve the reviewer's concern, we additionally conducted experiments on an alternative multi-view generative model V3D [3] to empirically demonstrate our framework's applicability on other diffusion models.
Due to the time constraint, we test this configuration on two cases.
The results (`v3d.mp4`) are included in the supplementary material.
Despite the slight Janus problem resulting from the limited capability of the adopted multi-view generative model, the generated 4D assets still exhibit overall spatial-temporal coherence, with significantly better consistency than that of the multi-view generative model-only baseline.
This result further demonstrates the potential of our approach to scaling with various foundational diffusion models.
**For Q3, the quantitative ablation for the decoupled reconstruction is not provided. I still think that limited qualitative results cannot fully support the effectiveness of different designs.**
Actually, there is a lack of comprehensive quantitative metrics to evaluate the quality of generated assets in this task, so we following the widely recognized works [1,2] by providing the qualitative ablations in our initial submission, which can clearly illustrate our advantage against the baselines.
In addition, we have also provided comprehensive quantitative ablations for the design of our first stage (which is the core part of the proposed framework) in our last response to Reviewer AvRx's Q3 (Quantitative results for ablation studies).
Now we additionally included the ablations on the decoupled reconstruction in Table 3 of the revised appendix.
From Table 3 (Appendix), we can observe that our full model (f) outperforms the baseline (g) without decoupled reconstruction.
**Moreover, the table results in Q3 are different from those inthe Appendix, please correct them.**
Thanks for the eagle eyes! Actually, the results in Q3 and those in the Appendix (Table 3) are essentially the same except for the format of CLIP score. The results in the response of Q3 were reported as percentages omitting the "%", while the numbers in the CLIP score column of Table 3 (Appendix) are presented in the actual decimal form. For example, 94.82 in Q3 and 0.948 in Table 3. We have fix it in the response of Q3.
**For L484 and L505, the titles for two paragraphs are same, please correct them.**
Thanks for pointing out this typo, we have fix it.
Hope our responses clarify the questions above. The reviewer’s constructive comments have been instrumental in guiding these refinements, and we are sincerely grateful for this guidance.
> [1] Tang, Jiaxiang, et al. Dreamgaussian: Generative gaussian splatting for efficient 3d content creation. *ICLR*, 2024.
> [2] Ren, Jiawei, et al. Dreamgaussian4d: Generative 4d gaussian splatting. *arXiv preprint*, 2023.
> [3] Chen, Zilong, et al. V3d: Video diffusion models are effective 3d generators. *arXiv preprint*, 2024.
> [4] Voleti, Vikram, et al. Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion. *ECCV*, 2024.