Response to Reviewer XVVX
Thank you for the response! In the following we provide additional explanations regarding to your questions. Please feel free to let us know if these address your concerns.
**Q1: It's recommended to plot the current dynamic coefficient curve.**
A1: Thank you for your suggestion. Unfortunately, **due to NeurIPS's requirement that no anonymous links be included in the rebuttal**, we regret that we cannot provide you with the dynamic coefficient curve at this time. However, we will certainly include this experiment and its analysis in the revised manuscript. Just a gentle reminder, **in Figure 10 of our paper**, we present the curve showing the gradient of the loss function (Eq 6) with respect to the coefficient during the denoising process. The figure illustrates that the coefficient exhibits significant convergence during denoising, and there are distinct update patterns for the coefficient when using different backbones in RealCompo.
**Q2: Is there exisits a general coefficient curve that is suitable for most of the prompt? If it is possible, there's no need for an optimization-based method that requires gradient backward, which will be mush easier for real application.**
A2: It is a valuable view to explore a general coefficient curve. However, we observed that the coefficient varies depending on the prompt based on extensive experiments. This is because **the coefficient is primarily optimized based on the layout, and different prompts correspond to different layouts, leading to varying coefficients**. We will certainly include visualizations and analysis of this part in the revised manuscript. Additionally, as seen in Figure 10, **RealCompo with different backbones also requires distinct coefficient update strategies**.
But I agree that having a general coefficient curve would make the application of our method more convenient and meaningful. Thank you for your thoughtful suggestions. This is a preliminary attempt, and we will explore the potential of a general coefficient curve in the future.
**Q3: It's recommended to show the results of the fixed coefficient (the best selected) instead of dynamic coefficient.**
A3: Thank you for your suggestion. **In the third column of Figure 8** (w/o Dynamic Balancer) in the paper, we present the results of experiments **using a simple fixed coefficient, where both models have the same coefficient**. The figure illustrates that without dynamically updating the coefficient, the T2I model, which lacks layout constraints, negatively impacts the positioning capability of the L2I model. This results in generated images where the object positions do not align with the layout, despite the T2I model retaining a higher realism advantage. We will include additional manually designed fixed coefficients in the manuscript to validate the effectiveness of our method.