HairDiffusion: Vivid Multi-Colored Hair Editing via Latent Diffusion

Hair editing is a critical image synthesis task that aims to edit hair color and hairstyle using text descriptions or reference images, while preserving irrelevant attributes (e.g., identity, background, cloth). Many existing methods are based on StyleGAN to address this task. However, due to the limited spatial distribution of StyleGAN, it struggles with multiple hair color editing and facial preservation. Considering the advancements in diffusion models, we utilize Latent Diffusion Models (LDMs) for hairstyle editing. Our approach introduces Multi-stage Hairstyle Blend (MHB), effectively separating control of hair color and hairstyle in diffusion latent space. Additionally, we train a warping module to align the hair color with the target region. To further enhance multi-color hairstyle editing, we fine-tuned a CLIP model using a multi-color hairstyle dataset. Our method not only tackles the complexity of multi-color hairstyles but also addresses the challenge of preserving original colors during diffusion editing. Extensive experiments showcase the superiority of our method in editing multi-color hairstyles while preserving facial attributes given textual descriptions and reference images.

Paper

References (46)

Scroll for more · 34 remaining

Similar papers

Peer review

Reviewer GkwL6/10 · confidence 2/52024-07-04

Summary

### NeurIPS Review #### Summary: This submission presents an innovative 2D hairstyle editing pipeline leveraging latent diffusion models (LDM). The key component of this method is the Multi-Stage Hairstyle Blend (MHB), which facilitates separate control over hairstyle and hair color. By integrating structural information from non-hair regions such as facial masks and keypoints, the method effectively preserves non-hair attributes in the editing results. Additionally, a hair warping module enables natural transfer of hair color from the source image to the generated target image. Extensive experiments on public datasets like CelebA-HQ demonstrate improved results compared to previous state-of-the-art methods. #### Strengths: #### Weaknesses:

Strengths

- Clarity and Writing: The paper is well-written and relatively easy to understand. - Technical Contributions: The submission presents a comprehensive system with solid and relatively novel technical contributions. * Hair Warping Module: The inclusion of a warping module that transfers the color pattern from the source hairstyle to the target hairstyle is notable. * Multi-Stage Hairstyle Blend (MHB) Module:** This module is effective in preserving non-hair attributes in the generated image. - Extensive Experiments: - The method produces more natural hair editing results compared to previous approaches, maintaining consistent hair color with the input text/image (control source) and better preserving non-hair regions. - A user study validates the human preference for the presented method over other methods. - The ablation study includes visualizations demonstrating the unique contributions of each module/control signal to the final results, along with detailed discussions.

Weaknesses

- More discussions on the failure cases: Most results are shown with a near-frontal head pose. Including and discussing more results with non-frontal head poses would enhance the completeness and understanding of the work. Additionally, as mentioned in the limitation section, the presented method might not work well with the color transfer between different hairstyles. It would also be great if more failure cases are included on that end. - Manuscript Quality: There are some typos and missing text. A thorough proofreading is necessary to refine the manuscript.

Questions

See weakness.

Rating

6

Confidence

2

Soundness

3

Presentation

2

Contribution

3

Limitations

Yes.

Authorsrebuttal2024-08-13

Due to the upcoming deadline for the discussion, I would like to confirm if you have any further questions or if there are areas where you need further clarification. If there are no more issues with my submission, I hope to receive your feedback and proceed. Thank you very much for taking the time to review my work amidst your busy schedule. I greatly value your opinions and hope to improve my research under your guidance. Thank you for your assistance, and I look forward to your reply.

Reviewer X6RQ4/10 · confidence 3/52024-07-11

Summary

This paper presents a new framework for hair editing tasks, which includes editing hair color and hairstyle using text descriptions, reference images, and stroke maps. The proposed approach leverages Latent Diffusion Models (LDMs) and introduces the Multi-stage Hairstyle Blend (MHB) technique to effectively separate the control of hair color and hairstyle. Additionally, the method incorporates a warping module to align hair color with the target region, enhancing multi-color hairstyle editing. The approach is evaluated through extensive experiments and user studies, demonstrating its superiority in editing multi-color hairstyles while preserving facial attributes.

Strengths

The proposed method demonstrates impressive performance in multi-colored hair editing, showcasing the ability to handle complex hair color structures while preserving facial attributes effectively. The integration of Latent Diffusion Models (LDMs) and the Multi-stage Hairstyle Blend (MHB) technique provides a novel approach to decoupling hair color and hairstyle, enhancing the quality of the edited images.

Weaknesses

1. The overall pipeline primarily consists of two stages: Altering the Hairstyle: This involves using a combination of control net and diffusion models with a hair-agnostic mask, 2D body pose keypoints, prompts, and a reference image as conditions. These modules are commonly found in previous methods. Editing the Hair Color: This stage blends information between the hair mask and source image, along with the Canny image of the stylized image and a prompt, to achieve the final result. The main contribution lies in the Multi-stage Hairstyle Blend (MHB) method, which warps the reference hair to the source image to better maintain the hairstyle. However, warping modules are also commonly used in previous methods [4, 21]. The overall technical contribution does not meet the bar for this conference. 2. The hair structure of the source image is not well-maintained during color editing. As demonstrated in Fig. 5, fourth row, the strand direction of the hairline in the generated image is different from the input image. The results of HairClipv2 better preserve the curliness structure of the original input hair. Additionally, the generated image in the third row of Fig. 5 shows noticeable brightness differences and a distinct color discrepancy compared to the input image. 3. The paper's writing lacks clarity. It would be beneficial to specify the difference between \( I_c \) and \( I_i \) in Fig. 2, within the caption of this figure. Additionally, using \( I_i \) to denote the style proxy creates confusion.

Questions

Please check Weaknesses.

Rating

4

Confidence

3

Soundness

3

Presentation

2

Contribution

3

Limitations

NA

Authorsrebuttal2024-08-13

Due to the upcoming deadline for the discussion, I would like to confirm if you have any further questions or if there are areas where you need further clarification. If there are no more issues with my submission, I hope to receive your feedback and proceed. Thank you very much for taking the time to review my work amidst your busy schedule. I greatly value your opinions and hope to improve my research under your guidance. Thank you for your assistance, and I look forward to your reply.

Reviewer Ubef6/10 · confidence 4/52024-07-11

Summary

The paper introduces an approach for hair editing using Latent Diffusion Models (LDMs). A warping module ensures precise alignment of the target hair mask and enables hair color structure editing using reference images. The proposed Multi-Hair Basis (MHB) method within LDMs decouples hair color and hairstyle. The authors showcase the performance of their method through GAN-based methods via qualitative and quantitative evaluations,

Strengths

The paper introduces a warping module ensuring precise alignment of the target hair mask. This helps in handling some small mismatches between the images. The method separates hair color from hairstyle. This allows flexibility and control over the hair transfer task. The paper also shows results of text-based hairstyle editing, reference image-based hair color editing, and claims to preserve facial attributes.

Weaknesses

1) Most of the examples shown in the results are front-facing subjects in the case of the hairstyle transfer. In real-world applications, there are cases where the reference and source images can have diverse/different poses. The lack of such results raises the question if this is a limitation of the method. Such problems are addressed in the paper HairNet[1]. 2) A limitation of some GAN-based approaches compared in the paper is that to achieve high-fidelity reconstruction, the images are overfitted into the networks. As such most of these methods do not support further edits, for example shortening the transferred hair, making it wavy or curly or changing the pose of the subject to view the subject from a different angle. The current method does not seem to solve this problem either. What are the advantages of these GAN frameworks? Some of these methods support further editing as well. (check video associated with HairNet). https://www.youtube.com/watch?v=WBB43cgCFZM&t=153s 3) Some of the GAN based approaches can also control the degree of hairstyle transfer (Barbershop). Does this method control this efficiently? [1] Zhu, Peihao, et al. "Hairnet: Hairstyle transfer with pose changes." European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022.

Questions

1) Most examples in the results are front-facing subjects. How does the method perform when the reference and source images have diverse or significantly different poses? 2) GAN-based approaches often overfit images for high-fidelity reconstruction, limiting further edits. How does the current method handle this issue? 3) Can your method support additional edits post-hairstyle transfer, such as shortening hair, making it wavy or curly, or changing the subject's pose to view from different angles? 4) What specific advantages does the current method offer over existing GAN frameworks, especially regarding editability and handling diverse poses? Also, check the weakness section.

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors have discussed the limitations.

Reviewer 8NSD4/10 · confidence 5/52024-07-12

Summary

This paper presents a novel approach called HairDiffusion for editing hair in images using latent diffusion models. The main contributions of the work are: 1. Introduction of the Multi-stage Hairstyle Blend (MHB) method for effectively separating control over hair color and hairstyle in the latent space of the diffusion model. MHB divides the diffusion process into two stages, allowing for precise guidance of hair color generation and context-aware generation of the rest of the image. 2. Development of a warping module to align hair color with the target area. This module adapts the HR-VITON architecture, using DensePose and segmentation maps to account for facial poses. 3. Utilization of hair-agnostic masks to transform the hair editing task into an inpainting task. The authors developed two types of masks for different editing stages, which effectively preserve necessary information and remove unnecessary information. 4. Fine-tuning of the CLIP model on a dataset with multi-colored hairstyles to improve editing of complex hair color structures. The authors applied data augmentation to increase pattern diversity. The authors demonstrate that their method outperforms existing approaches in hairstyle and hair color editing tasks, especially for complex multi-colored styles. The method allows for editing hairstyle and hair color separately or together, using textual descriptions or reference images. Experiments show that HairDiffusion better preserves relevant image attributes (e.g., facial features, background) compared to existing methods. The integration of these components into a single pipeline enables HairDiffusion to work effectively with both simple and complex multi-colored hairstyles while preserving other image attributes.

Strengths

A major strength of this method is the quality of complex hair color transfer, as well as the preservation of hair texture, which other methods typically don't pay much attention to. The approach demonstrates exceptional ability in handling multi-colored hairstyles. Additionally, the method's capacity to preserve relevant attributes such as facial features and background elements sets it apart from existing techniques. The flexibility to edit hairstyle and color separately or together, using either text or reference images, adds to its versatility and practical applicability.

Weaknesses

The paper exhibits several weaknesses. Firstly, it employs confusing notations, some of which are not properly introduced in the text. Secondly, the ablation study lacks clarity and sufficient detail. There is also a notable absence of comprehensive information regarding the method's limitations. For instance, the paper fails to provide examples demonstrating how the style transfer functions when transferring from long to short hair, how it handles cases with significant pose differences, or how it manages complex textures and hairstyles. Moreover, there is insufficient information about the limitations of the warping module. All presented images showcasing this module display identical hair shapes, suggesting that the warping module's functionality is not truly tested. This makes it impossible to accurately evaluate the quality of its performance. Furthermore, the paper lacks sufficient information for its reproduction, as it wasn't explained how exactly and on what data the CLIP model was fine-tuned. In general, there is very little information throughout the work about additional datasets, where and how the data was scraped, and about hyperparameters. The user study is very poorly described; there is no information anywhere about what exactly the questions looked like and in which domains these experiments were conducted. There are also questions about the statistical significance of the results. The scientific novelty of this work is questionable. As demonstrated in Figure 3, the Control-SD method produces hairstyles that are virtually identical to those generated by HairDiffusion. This similarity suggests that the main contribution of the work is specifically in the transfer of hair color rather than the creation of the hairstyle. The color transfer component primarily consists of two stage: the warping module (which is based on HR-VITON, published in 2022) and the Multi-stage Hairstyle Blend (MHB) method (which utilizes ControlNet and simple blending techniques). Given that these core components are largely derived from existing methods, the overall originality of the approach is limited. Overall, the paper suffers from a lack of comprehensive evaluation. It presents a limited range of visual comparisons and metrics, which hinders a comprehensive evaluation of the method's performance. Specifically, it lacks crucial metrics such as realism after editing (FID), realism with pose differences (FID), and full-image reconstruction metrics (PSNR, SSIM, LPIPS, FID, RMSE) as used in [1]. Moreover, the authors have not included comparisons with recent relevant works like StyleGANSalon [1] and HairFastGAN [2]. This omission of key comparisons and metrics significantly limits the reader's ability to fully assess the performance of the proposed method in relation to the current state-of-the-art approaches. [1] Sasikarn Khwanmuang, Pakkapon Phongthawee, Patsorn Sangkloy, Supasorn Suwajanakorn. StyleGAN Salon: Multi-View Latent Optimization for Pose-Invariant Hairstyle Transfer. arXiv preprint arXiv:2304.02744, 2023. [2] Maxim Nikolaev, Mikhail Kuznetsov, Dmitry Vetrov, Aibek Alanov. HairFastGAN: Realistic and Robust Hair Transfer with a Fast Encoder-Based Approach. arXiv preprint arXiv:2404.01094, 2024.

Questions

1. As a strong point of your method, you indicate excellent quality in preserving facial details, but you compare it with HairCLIP v2, which works in a relatively low-dimensional FS space in StyleGAN. Could you provide more examples of facial detail preservation on very complex images in comparison with methods like StyleGANSalon and HairFastGAN? 2. Could you provide more examples of the limitations of your method? We would like to see how the method works with large pose differences, complex textures, and very different hair shapes in the image domain. With these images, we would like to understand how well the concept preservation works in the CLIP space, how well the warp module works, and how well the entire method functions overall. 3. Could you provide information on how you collected the additional dataset with hair colors?

Rating

4

Confidence

5

Soundness

1

Presentation

2

Contribution

1

Limitations

The paper does not adequately address the limitations of the proposed method. The authors dedicate only two brief paragraphs to limitations, and these are relegated to the supplementary materials. This treatment is insufficient for a comprehensive understanding of the method's constraints. A more appropriate approach would be to include a detailed discussion of limitations in the main body of the paper. Furthermore, the authors should present visual examples demonstrating scenarios where each module, as well as the overall method, underperforms. These examples should be accompanied by in-depth analysis to provide insights into the reasons for these limitations and potential avenues for future improvements. Such transparency would significantly enhance the paper's scientific value and reproducibility.

Reviewer 8NSD2024-08-10

Thank you for your responses. However, there are still some concerns regarding the limitations of your method that were not fully addressed in the rebuttal. Regarding your answer A1, you demonstrate an improved reconstruction by referencing Table 2 and Figure 3. While this shows that your method outperforms in reconstruction tasks, it was already evident that Stable Diffusion's reconstruction is superior to StyleGAN's. However, Table 2 indicates that your method underperforms significantly compared to StyleGAN-based methods in single color transfer. It would be beneficial to explain why this occurs and provide examples to illustrate this point. Additionally, the impact of such transfers on facial details remains unclear and should be elaborated upon. Your answer A2 did not directly address my original question. I inquired about how your method performs when transferring hairstyles from the image domain and its general limitations. Instead, you focused on the limitations of your method when recoloring hair. In summary, the current presentation of the work lacks transparency, and there are several issues that need to be addressed. If the final paper is permitted to incorporate these important responses from the rebuttal, I believe it would be suitable for publication. Additionally, given the recent publication of the Stable-Hair paper [1], it would be valuable to include a comparison with this work in the future version of your paper. [1] Yuxuan Zhang, Qing Zhang, Yiren Song, Jiaming Liu, Stable-Hair: Real-World Hair Transfer via Diffusion Model, arXiv:2407.14078

Authorsrebuttal2024-08-12

The discussion the challenges and limitations of the current method in hairstyle transfer, color alignment, and retaining facial details, while suggesting future improvements and comparisons with other approaches.

Thank you for your suggestions, which have been very helpful in improving our paper. Discussion on the Performance Compared to StyleGAN-Based Methods 1. Reasons for Subpar Performance Compared to StyleGAN-Based Methods: - In cases where hairstyles differ significantly, our color alignment method, which is designed to generate colors more consistent with adjacent hues, may produce colors that deviate from the original hairstyle color. This discrepancy can lead to inconsistencies in the generated results. - The sparse matrix control provided by ControlNet's Canny map does not offer pixel-to-pixel control over hairstyle features. Alternatively, we could attempt to use a step-by-step approach that mixes multiple ControlNet models, allowing the model to focus on additional information. - When editing hairstyles based on reference images while retaining Diffusion text editing, setting the text to "hair" by default may introduce additional information that degrades the quality of the generated results. We could try separating the control of the text and ControlNet to balance the control over hairstyle structure and color in the future. We will include examples of these issues in future versions. Impact on Facial Details 2. Transfers and Facial Details: - As illustrated in Figure 3, StyleGAN-based methods show suboptimal retention of accessories (G, M, O) and unique facial features (I) when the training dataset contains few or no samples of these details. The preservation of multi-color hair (N, Q) is compromised because HairFastGAN's approach of decoupling hair color and hairstyle in the feature space makes it challenging to maintain the color structure. StyleGAN Salon employs post-processing enhancements that improve results for multi-color hair to some extent, but retaining hairstyle details remains challenging. Hairstyle preservation (A, B, C, E, T, S, R, N) suffers from the lack of masks to retain irrelevant information. Additionally, maintaining the background (P) and hand preservation (F, K, L, W) in cases where the hand obscures the hairstyle is also difficult without optimization. StyleGAN-based methods struggle to preserve detail for accessories or hands, particularly when the training dataset contains few examples of such features. - Conversely, the ControlNet Canny map, which provides line-based control, shows better performance in preserving hairstyle flow and individual hair strands. Discussion on Hairstyle Transfer Limitations 3. Limitations of Hairstyle Transfer: - Hairstyle transfer is not the primary focus of our work, but it is important to address the method's shortcomings. We use a straightforward approach of converting hairstyles into text vectors, which often fails to capture the complete details and local features of hairstyles, leading to noticeable discrepancies between the generated and original hairstyles. Recent studies suggest using encoders trained on face datasets to better capture facial details, which is a promising direction for future work. We will include additional figures and textual explanations to discuss these limitations in future versions. - We have also noted the paper on Stable-Hair, which differs from our approach by not decoupling hairstyle features, thereby preserving both hair color and overall hairstyle during the transfer, and not focusing on text-based editing of hairstyles. This paper has not yet released code or demos; we will conduct a comparison once these resources are available.

Authorsrebuttal2024-08-13

Due to the upcoming deadline for the discussion, I would like to confirm if you have any further questions or if there are areas where you need further clarification. If there are no more issues with my submission, I hope to receive your feedback and proceed. Thank you very much for taking the time to review my work amidst your busy schedule. I greatly value your opinions and hope to improve my research under your guidance. Thank you for your assistance, and I look forward to your reply.

Authorsrebuttal2024-08-13

Thank you for your valuable feedback. We followed a paper from [1](ECCV 2024) and internally collected an In-the-Wild dataset from social media platforms such as Instagram to evaluate the model's performance in multi-color hairstyle transfer. We ensure that these data are used solely for research purposes and have been approved by the IRB. Additionally, since our method does not require facial information but only hairstyle information, we processed these images to remove any personally identifiable information and confirmed that the data was obtained from public resources or with appropriate usage permissions. We will add a section in future versions detailing our specific methods for de-identifying sensitive information. **To address the potential unfair usage** of generating hairstyles for different groups, we propose the following solutions: 1. Diversity and Representative Datasets: When collecting datasets, including both testing and training stages, we strive to include samples from various races, genders, and age groups to minimize the risk of bias and discrimination. 2. Responsible Use and Ethical Guidelines: We emphasize the importance of responsible use and the establishment of clear ethical guidelines when utilizing these models, ensuring that their use aligns with ethical standards. 3. Output Detection: We have implemented a detector to identify and filter content that is not suitable for public viewing, such as explicit or violent images, thereby protecting the user's visual experience. 4. Model Generalization: We utilize the powerful generative capabilities of diffusion models to achieve generalization and diversity: - Generalization Capability: The diffusion model is designed to learn the universal patterns and features of various hairstyles, not just those of specific groups present in the training dataset. This generalization capability allows the model to recognize and generate hairstyles suitable for different individuals. - Generating Diversity: The diffusion model has the ability to generate novel and unseen data, which indicates that the model can not only replicate the hairstyles from the training data but also apply the learned features to generate hairstyles suitable for different groups. This diversity is achieved through the model's extensive coverage of data distribution during the generation process. This aspect will be discussed in future versions. [1] Choi Y, Kwak S, Lee K, et al. Improving diffusion models for virtual try-on[J]. arXiv preprint arXiv:2403.05139, 2024.

Authorsrebuttal2024-08-13

Thank you for pointing this out. We acknowledge that in our initial submission, we mistakenly selected the "NA" option for the IRB review requirement due to a misunderstanding of the scope of "research on human subjects." We apologize for this oversight. In the final version, we will correct this and clearly state that our user studies have been approved by the IRB. Regarding our use of internet images, we based our approach on the methodology described in [1](ECCV 2024). We internally collected an In-the-Wild dataset from social media platforms such as Instagram to evaluate and train our model's performance in multi-color hairstyle transfer. These data are used solely for research purposes and have been approved by the IRB. Additionally, since our method focuses only on hairstyle information and does not require facial recognition, we have processed these images to remove any personally identifiable information. We confirmed that the data was obtained from public resources or with appropriate usage permissions. We will add a section in future versions detailing our specific methods for de-identifying sensitive information and ensuring copyright compliance. We appreciate the feedback and will ensure that our data collection and processing methods align with ethical standards in future research. [1] Choi Y, Kwak S, Lee K, et al. Improving diffusion models for virtual try-on[J]. arXiv preprint arXiv:2403.05139, 2024.

Reviewer X6RQ2024-08-13

Thank you for the rebuttal and the detailed explanations provided. While I appreciate the clarifications, I still have concerns regarding the overall novelty of the work. Specifically, the points raised about 'Addressing the lack of paired datasets' and 'Preserving hairstyle structure' do not seem to significantly advance the state of the art. Moreover, in your rebuttal, you mentioned that 'color discrepancy outside the inpainting region is a common challenge in diffusion-based inpainting tasks.' However, in Figure 4, the claim that 'Our approach shows better preservation of irrelevant attributes' appears somewhat overstated given the context. This discrepancy raises questions about the extent of improvement your method offers. Given these concerns, I have decided to maintain my original rating. That said, I am open to reconsideration if other reviewers strongly advocate for the acceptance of this paper.

Authorsrebuttal2024-08-14

Thank you very much for taking the time to review my work amidst your busy schedule. **Innovativeness and Effectiveness** : Our approach has achieved notable results in the novel task of multi-color hair transfer. Previously, no method applied the warp model to hair color transfer. We believe that if existing methods in the hairstyle domain had utilized the warp module, it would be necessary to compare our approach with them and demonstrate that our method surpasses them in metrics. However, as this is our first introduction of this module, we have validated the effectiveness of our method through the following: - Ablation Study (pdf Table 1): We have conducted ablation studies to verify the effectiveness of each step and the integration of the warp module. - Evaluation Metrics (pdf Table 2): As shown in Table 2 of the supplementary materials, our method, based on a general hair editing self-transfer approach, outperforms previous methods in the task of multi-color hair transfer. While our metrics in the task of swapping monochrome hairstyles are not the best, they are comparable to earlier methods. Considering our focus on the multi-color task, we believe this is justified. **Preservation of Facial Details and Irrelevant Attributes:** We have not overstated our effectiveness in preserving irrelevant attributes. In fact, in most cases, our method excels in maintaining facial details, accessories, clothing, and background. Although there are some changes in skin tone on the left side of Figure 4, the retention of facial details and the background is excellent. For example, in Figure 5 of the paper, despite changes in skin tone, the clothes of the child behind are preserved much better than in the text2img comparison method. Therefore, from the overall quantitative metrics and the preservation of irrelevant attributes, our method is superior to the comparison methods. We are willing to provide additional image examples to further demonstrate our preservation of details in the background, clothing, and other aspects. I greatly value your opinions and hope to improve my research under your guidance. Thank you for your assistance.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC