Dear Reviewer,
Thank you for taking the time to read and review our paper! We address your comments in detail below.
> **I would advise the authors to provide clear insights through experiments and offer some specific suggestions.**
In response to your advice for clearer insights and specific suggestions through experiments, we have introduced a new table (Table 1) in the revised manuscript. This table is aimed at providing a straightforward guide for readers in choosing suitable algorithms for various fine-tuning scenarios. For instance, based on our comprehensive experiments, we suggest that LoHa may be more effective for fine-tuning multiple "easier" concepts, whereas LoKr with full dimension appears better suited for targeting single "harder" concepts. This insight is reflected in the revised conclusion of our paper to provide more specific guidance.
In the meantime, we want to emphasize that while our summarized findings and suggestions provide a useful overview, they should be considered with caution. The complexity of fine-tuning text-to-image models, as discussed in detail in appendix C, means that a simplified summary could hardly capture the full picture. In this regard, our detailed text discussion offers a more nuanced explanation of the various ways in which a hyperparameter can influence the results, and is indispensable for a more comprehensive understanding of the fine-tuning process.
> **I cannot evaluate this paper because I believe it is proper for a benchmarking and dataset track, not the main track.**
We appreciate your perspective regarding the suitability of our paper for a benchmarking and dataset track. While our work does include benchmarking, it is important to highlight that these are not its only contributions. A major part of our research introduces and elaborates on new fine-tuning method such as LoHa and LoKr. Moreover, the introduction of the library should also be regarded as an independent contribution. Generally speaking, our work aims to identify the specific fine-tuning methods that are more appropriate for a task of interest. Many works in this line of research (discussed in detail in Appendix A) have been previously accepted by the main track of the conferences, such as
- Zhiheng Liu et al. Cones: Concept neurons in diffusion models for customized generation. In ICML, 2023.
- Yuchao Gu et al. Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models. In NeurIPS, 2023.
- Zeju Qiu et al. Controlling text-to-image diffusion by orthogonal finetuning. In NeurIPS, 2023
- Nupur Kumari et al. Multi-concept customization of text-to-image diffusion. In CVPR, 2023.
- Nataniel Ruiz et al. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In CVPR, 2023.
Finally, we do not believe ICLR possesses a separate "dataset and benchmark" track. In fact, on the [call for paper page of ICLR 2024](https://iclr.cc/Conferences/2024/CallForPapers), datasets and benchmarks is clearly indicated as a valid subject area, even though in our case, as explained, benchmarking serves more to support and validate the novel contributions, rather than being the sole focus of our work.