ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling

We propose ID-to-3D, a method to generate identity- and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured in-the-wild image of a subject. The foundation of our approach is anchored in compositionality, alongside the use of task-specific 2D diffusion models as priors for optimization. First, we extend a foundational model with a lightweight expression-aware and ID-aware architecture, and create 2D priors for geometry and texture generation, via fine-tuning only 0.2% of its available training parameters. Then, we jointly leverage a neural parametric representation for the expressions of each subject and a multi-stage generation of highly detailed geometry and albedo texture. This combination of strong face identity embeddings and our neural representation enables accurate reconstruction of not only facial features but also accessories and hair and can be meshed to provide render-ready assets for gaming and telepresence. Our results achieve an unprecedented level of identity-consistent and high-quality texture and geometry generation, generalizing to a ``world'' of unseen 3D identities, without relying on large 3D captured datasets of human assets.

Paper

References (91)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer fD9E7/10 · confidence 5/52024-06-16

Summary

The paper introduces a novel method for generating 3D human heads guided by both identity and textual descriptions. The proposed method employs face image embeddings and textual descriptions to optimize a neural representation for each subject. By leveraging task-specific 2D diffusion models as priors and a neural parametric representation for expressions, the method achieves high-quality results while requiring minimal training. Extensive experiments demonstrate the proposed method could generate higher-quality 3D facial assets than state-of-the-art methods.

Strengths

1. The proposed method is novel and promising in the field of 3D facial asset generation. The use of task-specific 2D diffusion models as priors for 3D head generation is a novel technique that reduces the dependency on large 3D datasets. The method incorporates a neural parametric representation to disentangle expressions from the identity, which is a reasonable way to manage the complexity of facial dynamics and ensure identity consistency in generated 3D heads. 2. Experiments evaluation is strong and demonstrates the superiority of the proposed method over existing methods. The quantitative analysis and visualization are impressive. The visual results also clearly show the advantage over state-of-the-art methods. Detailed metrics and visualization highlight the method's ability to produce high-quality, identity-consistent 3D models, showcasing significant improvements in terms of details and textures. The evaluation covers various aspects such as generalization to new identities, handling of different ethnicities, and avoidance of oversmoothing, providing a thorough validation of the method's effectiveness. 3. The generated 3D models are of high quality, with detailed textures and geometry. The method ensures that the identity of the input image is well-preserved in the 3D output. The versatility of the method is also demonstrated, which provides practical utility beyond simple model generation.

Weaknesses

1. The generated 3D head models, while detailed, sometimes appear exaggerated and more like caricatures than realistic human faces. This diminishes the method's applicability in scenarios where photorealism is critical. Addressing this issue would involve refining the model to balance between preserving identity and achieving realistic facial features. 2. The performance and quality of the method on significantly larger and more complex datasets are not extensively discussed. Evaluating and demonstrating the scalability of the approach would enhance its credibility and applicability. Discussing potential limitations when scaling up and providing strategies to handle larger datasets would be valuable additions. It would be interesting to discuss or show what could be further improved if larger datasets are incorporated. 3. The writing can be improved for better clarity and accessibility. For instance, it should be explicitly described that each identity requires a training stage. This would help readers understand the necessity and implications of the training process. Additionally, simplifying and clarifying the description of the multi-stage pipeline would make the methodology more accessible to other researchers and practitioners.

Questions

See the weaknesses section.

Rating

7

Confidence

5

Soundness

4

Presentation

3

Contribution

3

Limitations

Limitations are discussed in the paper.

Reviewer G3ux6/10 · confidence 4/52024-07-10

Summary

The paper presents a novel technique for creating 3D human head models from a single real-world image, guided by identity and text. This method is based on compositionality and uses task-specific 2D diffusion models as optimization priors. The authors extend a base model and fine-tune only a small training parameters to create 2D priors for geometry and texture generation. Additionally, the method utilizes a neural parametric representation for expressions, allowing the creation of highly detailed geometry and albedo textures.

Strengths

1. The method demonstrates significant innovation in the field of 3D head generation, particularly in generating high-quality models without the need for large-scale 3D datasets. 2. The capability to generate 3D models with disentangled expressions is a notable advancement, as this has been a challenge in previous research.

Weaknesses

1. The paper should discuess dataset diversity. Does the training dataset used in this study possess enough diversity to prevent potential biases? 2. There's a lack of analysis on the robustness of ID embeddings. Are these identity embeddings robust enough to accurately depict identity features across various expressions and poses? 3. Some citations are missing, specifically at line 191 where template-based approaches require references. 4. Is it logical for geometry diffusion to be fine-tuned using style-transfer methods? It could be reasonable if a pre-trained standard SD model like RichDreamer[1] is utilized. 5. Some results aren't convincing; for instance, in the last row of Fig9 within supplementary materials, reconstruction seems to lose eye-catching hair details. 6. Supplementary materials have been placed in a separate zip file instead of being attached to the main paper, which might breach some submission rules. [1] RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D

Questions

* Discussing the diversity of the dataset during training. * Discussing the robustness of id embeddings * Correcting citations. * Explaining why fine-tune using a geometry diffusion model with style-transfer instead of another geometry diffusion model.

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors acknowledged the limitations and deliberated on the possible negative impacts of the suggested technology on society.

Reviewer vm8B5/10 · confidence 3/52024-07-11

Summary

This work proposes a new approach for the generation of 3D human heads, which enables guidance with identity, facial expressions, and text descriptions. The approach is structured around two principal components: 1) the authors fine-tune a previously established text-to-image diffusion model through LORA on a specialized dataset of 3D human head models, to obtain the 2D guidance with separated texture and geometric details. 2) the method executes the generation of geometry and texture in separate stages, utilizing the SDS loss to optimize the process. It considers specific designs, such as the learnable latent code of facial expression, to enrich the geometry and texture details. It obtains better performance than several existing works.

Strengths

1. The proposed method is reasonable. The finetuning of the diffusion model on a specific 3D head dataset provides better guidance for geometry and texture. 2. It supports various types of conditional inputs, including identity, expression, and text description, thereby enriching its versatility in application. 3. It obtains better performance than several existing works.

Weaknesses

1. The generated head is of low visual quality. We are aware that text-to-image diffusion models are capable of producing very high-resolution images. However, the learned facial textures (resolution, clarity) showcased in this study are relatively poor. In comparison, other single-image 3D reconstruction methods, such as 3D GAN inversion combined with technologies like NeRF, can achieve very high visual quality. The text guided 3D portrait generation method, Portrait3D (siggraph 24), is also of high visual quality. 2. The innovativeness of the method is modest. Fine-tuning diffusion models on specific datasets and using SDS as a supervisory loss for 3D modeling are both fairly common practices. The methodological innovation in this study seems insufficient. 3. The expressions and text-guided editing scenarios demonstrated in the experimental section are quite basic (such as "eyes closed," "brow lowerer," "de-aged"), which limits their practicality. It is suggested to showcase more practical editing effects, such as changes in hairstyle or face shape, richer text-based guidance, to better understand its editing performance.

Questions

There are a few questions that should be addressed in the rebuttal, please see paper weakness for more information. ******** after rebuttal Thanks for providing additional experiments, that addressed some of the concerns. I raised my rating to borderline accept, mainly for its contribution of providing a relatively complete method for simultaneously controlling facial ID, expressions, and characteristics. Yet, the generated head, especially the texture, is still low in quality. Although this may be influenced by the dataset used, it is still a limitation of the method as there is currently no better dataset available (per the author's rebuttal), and it is also unlikely to construct a higher-quality dataset. Is there any other possible ways to further improve the image quality?

Rating

5

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

yes

Reviewer vm8B2024-08-13

Will the code and models be open sourced?

Will the code and models be open sourced?

Authorsrebuttal2024-08-13

Code publicly available

Yes, we confirm that the code and the models will be made publicly available.

Reviewer HnyJ5/10 · confidence 4/52024-07-12

Summary

**Summary This paper focuses on the task of 3D head generation. Specifically, the authors first extend a traditional diffusion model to a text-to-normal version and a text-to-albedo version with ID-aware and expression-aware cross-attention layers. Then, with the trained diffusion models, the authors optimize a neural parametric head model with a score distillation sampling loss. Extensive experiments demonstrate that the proposed method outperforms existing text-to-3D and image-to-3D methods in terms of 3D head generation. However, I have some concerns about this paper. My detailed comments are as follows.

Strengths

**Positive points 1. The authors introduce the first method for arcface-conditioned generation of 3D heads with score distillation sampling loss. 2. The proposed method can also achieve ID-conditioned text-based 3D head editing (e.g., age editing, changing hair color and gender).

Weaknesses

1. Although the proposed method generates a similar geometry to the input identity, the synthesized texture appears much unrealistic. What might be the cause of this phenomenon? Additionally, some implicit 3D representations, like NeRF [A-C] and 3DGS [D-F], can model high-fidelity surfaces for 3D heads. Why do the authors choose DMTET over these representations? It would be better if the authors could provide more discussion about the above questions. 2. The authors use five images as identity references for each 3D head. How is this optimal number determined? What is the relationship between identity similarity and the number of reference images? Quantitative results in terms of this should be provided. 3. There are some misalignments between Table 1, Figure 3, and Figure 4. For example, Fantasia3D is included only in Table 1 but does not appear in Figure 3 or the right column of Figure 4.. 4. In the original paper of Fantasia3D [G], the proposed method can only synthesize 3D assets given a text prompt as input. How do the authors adjust this methods to generate 3D heads when conditioned on specific identity? 5. For each 3D head asset, the authors extract identity features from multiple RGB images. How are these features combined? Are they added or concatenated? It would be helpful to provide more details about this operation. **Minor issues 1. On page 7, line 263, there is a missing space between “ID-to-3D” and “as”. **Reference [A] NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. ECCV 2020. [B] Implicit and Disentangled Face Lighting Representation Leveraging Generative Prior in Neural Radiance Fields. TOG 2023. [C] Geometry-enhanced Novel View Synthesis from Single-View Images. CVPR 2024 [D] 3D Gaussian Splatting for Real-Time Radiance Field Rendering. SIGGRAPH 2023. [E] Photorealistic Head Avatars with Rigged 3D Gaussians. CVPR 2024. [F] Relightable Gaussian Codec Avatars. CVPR 2024. [G] Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation. ICCV 2023.

Questions

Please refer to the weakness section.

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Yes.

Reviewer HnyJ2024-08-13

The authors have addressed my concerns. I have decided to maintain my score as a Borderline accept.

Reviewer fD9E2024-08-13

Maintain rating

The rebuttal addressed some of my concerns. I recommend the authors explain the exaggerated caricature-like style in the paper and maybe need to change the "realistic" claims. I would maintain my initial rating.

Authorsrebuttal2024-08-13

We thank the reviewer for the feedback. We appreciate the suggestions and will ensure to comment on the caricature-like style in the limitations section of the paper and revise the claims as suggested.

Area Chair eNJE2024-08-13

Response to Authors

Dear Reviewer G3ux, The authors have responded to the questions that you raised in your review. The discussion period with the authors is soon coming to a close, at 11:59 PM AOE Aug 13, 2024. There is still time to respond to the authors. It would be great if you could acknowledge that you have read the authors' response, engage in discussion with them and update your final score. Best, AC

Reviewer G3ux2024-08-14

Thanks for the author’s response. I would raise the score to 6.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC