DreamHuman: Animatable 3D Avatars from Text

We present DreamHuman, a method to generate realistic animatable 3D human avatar models solely from textual descriptions. Recent text-to-3D methods have made considerable strides in generation, but are still lacking in important aspects. Control and often spatial resolution remain limited, existing methods produce fixed rather than animated 3D human models, and anthropometric consistency for complex structures like people remains a challenge. DreamHuman connects large text-to-image synthesis models, neural radiance fields, and statistical human body models in a novel modeling and optimization framework. This makes it possible to generate dynamic 3D human avatars with high-quality textures and learned, instance-specific, surface deformations. We demonstrate that our method is capable to generate a wide variety of animatable, realistic 3D human models from text. Our 3D models have diverse appearance, clothing, skin tones and body shapes, and significantly outperform both generic text-to-3D approaches and previous text-based 3D avatar generators in visual fidelity. For more results and animations please check our website at https://dream-human.github.io.

Paper

Similar papers

Peer review

Reviewer FnWG6/10 · confidence 4/52023-06-27

Summary

This is a paper focusing on generating 3D animatable full-body human avatar from text using pretrained 2D diffusion model and Score Distillation Sampling (SDS). The proposed approach differs from the existing approach in that, instead of directly representing a canonical space using surface template e.g. SMPL, it 1) uses the implicit 3D human model to establish the correspondence and condition the canonical representation on the pose parameter, 2) adopts a per-part optimization, and 3) uses a physics based shading formulation and jointly optimizes the environment lighting. These three technical novelty intent to achieve more plausible deformation for loose clothing, better detail reconstruction for faces and hands, and more realistic colors, respectively. The comparisons with AvatarClip and DreamFusion demonstrate the advantages of the proposed method.

Strengths

- substantially better visual quality compared to AvatarClip and DreamFusion, although the former only with limited evidence. - the use of imGHUM seems to improve the diversity of clothing. - the part-based optimization visibly improves the visual quality of the generation.

Weaknesses

- Even though the following papers can be considered concurrent, they should be mentioned in the related work. Cao, Yukang, et al. "Dreamavatar: Text-and-shape guided 3d human avatar generation via diffusion models." arXiv preprint arXiv:2304.00916 (2023). Jiang, Ruixiang, et al. "AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control." arXiv preprint arXiv:2303.17606 (2023). - The comparison with AvatarClip is not sufficient. In the supplemental material and Tab 1, results of AvatarClip should be included. - Since the NeRF in canonical space is pose-dependent, it is theoretically more prone to overfitting. How does the method perform for unseen poses? - The fact that the shape parameter $\beta$ can vary is not well explained. - The benefit of the shading and optimization of the SH for environment light is not elaborated sufficiently. What is the albedo before shading? How exactly do you model the irradiance (include rendering equation). The training trick with randomly perturbed SH coefficients is not well motivated. If the goal is to improve disentanglement, why not use some smoothness regularization? - The animations are shown in a fixed view. It's hard to judge whether loose clothing such as dresses and jackets deform as claimed from a fixed viewing angle.

Questions

It would be great to address my comments above. In particular, I look forward to seeing in the rebuttal: 1. Comparison with AvatarClip 2. Clarification about shading and visualization of the learned albedo 3. Some visual examples of some avatars with loose clothes in walking motion, shown from the frontal view. 4. It would be great to include a user study on the visual quality.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

2 fair

Presentation

3 good

Contribution

3 good

Limitations

Yes

Reviewer hMhh6/10 · confidence 4/52023-07-01

Summary

This paper presents a method to generate animatable 3D human avatars from text. The pipeline is similar to DreamFusion and is built upon the Nerf representation and diffusion model. However, a key difference is that an imGHUM body model is introduced as prior, which allows for the construction of a deformable Nerf representation and a 3D animatable human model. This design not only enables animation capabilities but also effectively addresses the anthropometric consistency issues. In addition, a semantic zooming loss is proposed to refine details in body regions like the face and hands, resulting in a more photo-realistic overall quality. Quantitative and qualitative comparisons are conducted with state-of-the-art baselines such as DreamFusion and AvatarCLIP. The extensive results demonstrate that the proposed method outperforms previous approaches across all metrics. The visual quality and geometry detail are particularly impressive. Furthermore, the study includes an analysis of different components to highlight their importance within the framework.

Strengths

- The paper is well-written and easy to follow. - The overall results are impressive, especially the appearance and the geometry details of the generated 3D human avatar. - Although each component can be seen in previous works like DreamFusion and AvatarCLIP, this work did a good job on putting all losses and modules together properly and achieving promising 3D human avatar modeling. - Extensive ablation experiments are conducted to show the importance of proposed components. - The proposed semantic zooming loss is interesting and effective, which largely improves the visual quality and helps to generate sharper, higher-quality textures.

Weaknesses

- It’s not clear to me how to decide the shape parameters for the imGHUM model during optimization. If it is optimized together with the Nerf model, will it introduce additional training costs? - The overall computation cost is not clearly listed and compared. It would be better to report the optimization time and inference time for a single model/text prompt for the proposed method and baselines. - It would be better to show more quantitative results for the semantic zooming loss since it's one of the key contributions of this work.

Questions

- Is there any quantitative result for more direct evaluation regarding the view consistency and pose dependency of the generated 3D avatar?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

- As shown in Figure 7, artifacts can be found in the soldier avatar generated by DreamHuman. The generated avatar has more than two hands, and the method struggles with generating correct accessories. This can be really hard for the current model design since it's built on imGHUM body model.

Reviewer hMhh2023-08-16

Thanks for your response.

Thank you for the detailed feedback and the new qualitative results, which address my earlier concerns. I believe this is solid work with several innovations. And I agree that making the model accessible for the purpose of reproducibility would serve as valuable assets for the broader community.

Reviewer urhH8/10 · confidence 5/52023-07-05

Summary

This work proposes a method for text-driven human avatar generation. It combines animatable human nerf and diffusion model to implement avatar generation and animation. This work produces photorealistic avatars with high-quality details by incorporating spherical harmonics lighting model and semantic zoom. Extensive experiments demonstrate SOTA performance and the effectiveness of each design in the proposed framework.

Strengths

1. This is the first diffusion-based work that successfully produces photorealistic animatable 3D human avatars. 2. This work shows temporally consistent animation results. I believe this can open up more application possibilities for optimization-based avatar generation methods. 3. The incorporation of spherical harmonics lighting model can alleviate the long-standing issue of unrealistic over-saturated color for text-driven 3D object generation. 4. The semantic zoom loss is simple yet effective to improve the quality for detail regions such as face, arm, and hand.

Weaknesses

1. The imGHUM is designed for the whole body, it contains parameters for hand and facial expressions. But there is no result for the animation of facial expressions and hand poses. It could be better if the authors can provide animation results to show the controllability of these details. I think this will also help prove the necessity of semantic zoom loss. 2. I believe this work uses a more powerful diffusion model, and imGHUM is not completely open source. These two points will limit access to the proposed model. Is there any plan to release an online demo or interface for users? 3. Although the authors mentioned the training strategy in Sup Mat, there is no clear description of the computation cost of the proposed method.

Questions

1. How to determine the rendering camera poses for each part in the semantic zoom? 2. Which diffusion model is used in this work? Is this diffusion model finetuned on human body images?

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

Yes, addressed.

Reviewer KMW96/10 · confidence 3/52023-07-09

Summary

This paper proposes a method to generate high-quality and animatable 3D human from textual input.

Strengths

- The result is good, showing clear improvement compared to the previous text-to-3D method. - The method can be learned without 3D GT. - The ablation study is thoroughly done.

Weaknesses

- What is the run-time at inference? - Density loss: I think this can work when the suject is wearing a tight clothing. However, wouldn't this confuse the network when the subject is wearing a clothing with large deformatation (i.e., more gap is present between the body model and the actual geometry)? - What is w_i of equation 5? - What is the proposal weights L_p in L195? - It is written that 4 renderings are supervised for a single training step (supplementary material L7). Does this mean that 4 randomly selected semantic parts (zoomed-in semantic parts) of the same pose are rendered? Or does this mean that single semantic part is rendered from different camera pose and body pose? Also, would rendering and supervising predictions less than four (e.g., when using single GPU with less memory) degrade the performance? Discussion on the running environment and performance would be helpful for the readers. - Lack of implementation details, making it difficult to reproduce. Also, no plan on supporting the reproducibility has been presented.

Questions

- What is the rendering resolution at inference time? - How are the query locations sampled? Are they sampled in free space or bounded box? - Which method was used to perform rendering with spherical harmonics lighting model? - How are the camera parameters (both extrinsic and intrinsic) set? Since there are no known camera parameters for the images generated with Diffusion models, I am wondering how the authors set the parameters to render images. - Is there a reason behind choosing mip-NeRF360 as the backbone? Would using basic NeRF degrade the performance? - What is s (the output of imGHUM) exactly? Is it an index of the nearest vertex on the body? - Although the lighting result is much better than the previous work (Dream Fusion), still it looks unnatural. What is the reason behind this and how it can be improved?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

- The idea is interesting and the results are great. However, there is a concern regarding the reproducibility.

Reviewer urhH2023-08-18

Thank the authors for the answers. The results in the rebuttal file are nice and persuasive. All of my concerns have been addressed, and I will keep my rating.

Reviewer FnWG2023-08-18

Satisfied with the answers

Thank you for the explanation. I think the answers to shape parameter variation and optimizing environment light SH is important details, and should be included in the main paper. I appreciate the comparison with concurrent work. Finally, I agree with other reviewers in urging the authors to release their code. My final rating is accept.

Reviewer KMW92023-08-18

Response to the author rebuttal

I appreciate the authors for their time and efforts put into the rebuttal. The author rebuttal has successfully addressed my concerns. I will keep my rating.

Program Chairsdecision2023-09-21

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC