Thank you for taking the time to proactively participate in the author-reviewer discussion. We would like to clarify your concerns raised in the latest comment.
__Q1__ -> Since we are unable to provide turn-table videos (which would have been highly efficient in showcasing the anatomic inconsistencies) in the rebuttal, we would humbly request the reviewer to observe the generated outputs of LucidDreamer in Fig. (a). It can be noted that in the front-view, __the elephant's tusk is visible in between the front legs and it goes all the way down to the bottom of the legs. While in the side-view the trunk is shown as raised in the air. In the front-view the elephant's front legs are shown as straight and touching the ground, but in the side-view, upon zooming you can identify that there is a bent leg and two straight legs on the front side of the animal. The generated asset also has two tails that are visible upon closer inspection in the side-view. We would also like to highlight that it is highly green.__ For the “llama with octopus tentacles body”, __LucidDreamer generates a full-body llama with a single tentacle around the head__, but with YouDream we were able to explicitly guide the generation of an animal with the llama head and tentacles as legs, where even the number of such tentacle-legs is controlled accurately using the 3d pose input.
We would also like to discuss other inconsistencies we observed in other assets we generated using LucidDreamer which we could not include due to space, such as “a giraffe, full body” only produced the head and neck of the giraffe, and “a three-headed dragon, full body” had a dragon with a single head with wings protruding from the head.
Due to time constraints of the rebuttal period, we were able to test 3 new methods – Stable Dreamfusion, ProlificDreamer, and LucidDreamer and were not able to test additional methods such as GaussianDreamer, Magic123. However, since GaussianDreamer does not provide any explicit control/signal to avoid the generation of anatomical inconsistencies, such as multiple legs, heads, or tails, we believe it will contain similar issues as shown in LucidDreamer, HiFA, ProlificDreamer, etc. which are guided by text-to-image diffusion models.
We do agree that the visual fidelity of some of the prior methods, such as MVDream and LucidDreamer (uses SD 2.1 while YouDream uses SD1.5), is higher owing to certain factors such as 3D training data, higher NeRF dimensions, higher SD versions, and various other tricks shown in the prior art. Hence, __in this paper, we focus on improving the anatomical consistencies which can be very easily alleviated by the pose conditioning offered by YouDream__. To showcase the impact of such tricks on visual fidelity we also added results using slightly higher dimension NeRF in Fig. (c) (due to resource constraints we could only increase the NeRF dimensions by 2x).
__Q2__ -> During the rebuttal period, we were able to produce 10 assets of LucidDreamer – “a tiger, full body”, “a giraffe, full body”, “a three-headed dragon, full body”, “a pangolin, full body”, “an elephant, full body”, “a llama with octopus tentacles body, full body”, “a giraffe with dragon wings, full body”, “a realistic mythical bird with two pairs of wings and two long thin lion-like tails, full body”, “a red male northern cardinal flying with wings spread out, full body”, and “a Tyrannosaurus rex, full body”. Only using these assets for a CLIP-score-based evaluation of LucidDreamer, MVDream, and YouDream, we obtained the following scores:
_____________________________________
_LucidDreamer_ | _MVDream_ | _YouDream_ |
_____________________________________
28.89 | 29.13 | 30.29 |
______________________________________
These numbers are using the same setting as described in the paper. Please note that in the paper we compare using 22 generated assets. We will extend this evaluation to all 22 assets for our final version.
We would also like to justify the reason behind not comparing CLIP-Score with prior text-to-image-based 3D generators and only comparing with MVDream, which was trained on 3D data. Since the core contribution of the paper is to enable the generation of anatomically consistent animals, we found that every text-to-image-based model produced some or the other geometrical artifacts resulting in inconsistent 3D animals. Hence we only compare with MVDream which can recreate the geometric consistencies of natural animals (but fails for unseen animals) thanks to its 3D (multi-view consistent) pre-training.