Debiasing Scores and Prompts of 2D Diffusion for View-consistent Text-to-3D Generation

Existing score-distilling text-to-3D generation techniques, despite their considerable promise, often encounter the view inconsistency problem. One of the most notable issues is the Janus problem, where the most canonical view of an object (\textit{e.g}., face or head) appears in other views. In this work, we explore existing frameworks for score-distilling text-to-3D generation and identify the main causes of the view inconsistency problem -- the embedded bias of 2D diffusion models. Based on these findings, we propose two approaches to debias the score-distillation frameworks for view-consistent text-to-3D generation. Our first approach, called score debiasing, involves cutting off the score estimated by 2D diffusion models and gradually increasing the truncation value throughout the optimization process. Our second approach, called prompt debiasing, identifies conflicting words between user prompts and view prompts using a language model, and adjusts the discrepancy between view prompts and the viewing direction of an object. Our experimental results show that our methods improve the realism of the generated 3D objects by significantly reducing artifacts and achieve a good trade-off between faithfulness to the 2D diffusion models and 3D consistency with little overhead. Our project page is available at~\url{https://susunghong.github.io/Debiased-Score-Distillation-Sampling/}.

Paper

Similar papers

Peer review

Reviewer nz3q5/10 · confidence 5/52023-06-10

Summary

The paper proposes two simple methods to solve the widely known Janus problem in zero-shot text-to-3D generation, which is an essential issue. The proposed methods are intuitive and simple. They are more like optimization tricks instead of sysmetic formulations. From the qualitative results, the improvement of the proposed methods is not obvious.

Strengths

1)The motivation is persuasive because the Janus problem is very important in text-to-3D generation. 2)The paper proposes two novel strategies of debiasing the score-distillation from 2D diffusion models, solving the Janus problem widely existing in zero-shot text-to-3D generation. 3)The first strategy performs dynamic clipping of 2D-to-3D scores to eliminate the biases towards some viewing directions, which solves the problems of additional legs, beaks, and horns. 4)The second strategy uses a pre-trained LLM to identify and remove the conflict words with the view points.

Weaknesses

1)Technically, the novelty is a little weak. 2)The paper proposes to debias scores and prompts of 2D text-image diffusion models. But these two debiasing methods are simply integrated for text-3D generation without elegant co-formulation. 3)Even with sophisticated mathematical formulations, the actual debiasing score method is very intuitive and simple. 4)The prompt debiasing method is also very simple and intuitive. What is the difference between the proposed method and using the CLIP similarity scores to remove conflict words? 5)The qualitative comparisons in Figure 6 and 7 do not show much superiority of the proposed method.

Questions

See weakness

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

See weakness

Reviewer 9gp36/10 · confidence 5/52023-06-28

Summary

This paper proposes two approaches to debias the score-distillation frameworks for view-consistent text-to-3D generation. The first approach is called score debiasing, involves cutting off the score estimated by 2D diffusion models and gradually increasing the truncation value throughout the optimization process. The second approach, called prompt debiasing, identifies conflicting words be tween user prompts and view prompts using a language model, and adjusts the discrepancy between view prompts and the viewing direction of an object. The proposed results is demonstrated by experiments.

Strengths

1. This paper is clearly written. 2. The proposed method is technically sound and insightful. 3. Some experiments are conducted to demonstrate the idea.

Weaknesses

1. The experiment is not sufficient, as discussed in Questions.

Questions

1. My main concern is about the evaluation. Since the experiments are only conducted with 70 prompts, could you report the success rate of generation, i.e., how many generated results of the 70 prompts does not have Janus problem? This may be more convincing than the metrics used in the paper. 2. I have tried the score debiasing method in Stable-DreamFusion repo. However, based on my observation, the Janus problems are not alleviated. Is it true that the improvements is actually from Prompt Debiasing, and the improvement of Score Debiasing is quite marginal? If it is not true, could you provide an experiment solely on Score Debiasing and report the improvements of success rate of generation? (Success = Without Janus problem, Fail = With Janus problem.) I am glad to raise my score if the concerns are resolved.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

1. Since the results are based on NeRF with latent space rendering, the quality of generated 3D results is lower than current SOTA. Could this method be applied to other 3D representations, like DMTet, which demonstrates better quality than NeRF as is shown in Magic3D? Could this method be applied to a NeRF that renders in rgb space which demontrates stronger 3D consistency?

Reviewer gQ7C6/10 · confidence 3/52023-07-05

Summary

This paper addresses the Janus problem appearing in Score-Distillation-based 3D generation methods, where the most canonical view of an object appears in other views. In particular, the author propose two components for debiasing the score distillation and the prompt used for the generation. First, the authors provide a theoretical discussion over SDS loss and the terms that contribute to the Janus problem and bad generation quality. Therefore, the authors propose gradient clipping of the score gradient in order to tackle the discussed problem. Furthermore, the paper discusses another issue of the existing methods, which is the potential contradiction between tokens of the text prompt with the added view guidance tokens. To prevent the contradiction, the authors propose to exploit the point-wise mutual information (PMI) technique to identify and remove the contradictions from the text prompt. The proposed method evaluated quantitatively and qualitatively and compared to existing baselines. In addition, an ablation study on different proposed components is provided.

Strengths

-- The paper is well-written. -- The provided discussion on the bias of score distillation loss and view-augmented prompts is valuable -- The novelty of the proposed components is sufficient. -- Based on the provided results, the proposed method seems effective in alleviating the Janus problem.

Weaknesses

-- The efficacy of the proposed method is still highly limited by the bias of diffusion models. (although this is an issue specific to the proposed method).

Questions

-- I am wondering how much the proposed gradient clipping affects the convergence speed of the optimization. Is there any considerable difference with the baselines regarding the optimization steps?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Limitations and the societal impact of the work has been discussed in the supplementary material.

Reviewer 8eVk6/10 · confidence 3/52023-07-06

Summary

This paper proposes debiased score sampling (D-SDS) for improving 2D diffusion-based 3D generation, targeting at the Janus problem. The method is composed of two parts: (i) score debiasing that cuts off scores from 2D diffusion and (ii) prompt debiasing that fixes the discrepancy between view prompts and object orientation. Experiments have been conducted from the baseline method SJC [Wang et al., 2023], where impressive improvements have been shown, especially in solving the Janus problem. A detailed ablation study has shown the value of each debiasing design. [Wang et al., 2023] Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D Generation. In CVPR.

Strengths

- This work is well-motivated, targeting a relevant problem of the Janus problem in 3D generation. The discussions on the problem and technical derivations are solid, insightful, and interesting. Besides, the paper is well-written. - The proposed method is novel. I like the idea of debiasing scores and prompts, which makes good sense to me. Besides, the proposed method can be easily extended to other 2D diffusion-based methods. - Experiments are well-designed and presented. Qualitative, quantitative, and user study experiments are shown, and the results demonstrate the effectiveness of solving the Janus problem.

Weaknesses

- Somewhat not essential. The proposed method is simple and effective. However, it seems the method is only applicable to these SJC/DreamFusion 2D diffusion score-based methods. This is not a fundamental solution for the Janus problem but is kind of like a temporary effective trick that alleviates problems. - Lacking complex examples. The presented results use relative prompts mainly for simple 3D objects. What about more complex 3D structures like the "Temple of Heaven" shown in SJC? Since one of the main difficulties of text-to-3D generation is for these more comprehensive geometries, it is important to show more of these cases. Is it another limitation of the proposed method? - Limited baseline. The proposed method is currently conducted to improve the SJC baseline. What about other methods? Is it possible to be applied to the more recent work ProlificDreamer [Wang et al., 2023]? - Minor suggestion: It seems that the Janus problem in 3D generation was first pointed out and termed for DreamFusion by Ben Poole before on [social media](https://twitter.com/poolio/status/1578045212236034048?s=20), it would be better to cite DreamFusion when first listing this issue in the paper. But it is okay if it is not cited for this term. [Wang et al., 2023] ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation. arXiv preprint. [Poole et al., 2023] DreamFusion: Text-to-3D using 2D Diffusion. In ICLR.

Questions

I have no other questions besides those listed before. I am happy to increase the score if my concerns are addressed.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Since the method is based on SJC, it will strongly bound its performance.

Reviewer 9gp32023-08-11

I have read your rebuttal and it partially solves my concerns. I have raised my score to "weak accept". I suggest that you add the table of success rate and the experiments in the attached pdf into your paper somewhere in the main text or appendix, if accepted.

Authorsrebuttal2023-08-13

Response

Thank you for your feedback and for considering updating the scoring for the paper. I appreciate your suggestion and will revise our paper accordingly.

Reviewer gQ7C2023-08-11

Final Comment

I thank the authors for their response and answering my question. I keep my initial score.

Reviewer 8eVk2023-08-12

Post Rebuttal Comment

I highly appreciate the authors for the detailed response, which I suggest adding to the main paper appendix in the final version. My concerns are largely solved. Minor suggestion: It would be better to include some quantitative comparisons for methods like ProlificDreamer and Magic3D, if possible.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC