Summary
The manuscript introduces propose the strategies to resolve the limitation of overfitting to the small subset of training images and hard to generalize to new contexts with different text prompt input. In practice, the method aims to create a rich set of text prompt and the use of unsupervised learning objective to enhance the ability of concept personalization with the new text prompt input. The results demonstrate that the method is competitive with leading frameworks in various image generation tasks.
Strengths
The authors have developed a framework for generating personalized images that effectively integrates the context diversification of personal concept using masked language modeling and to solve the issue of concept overfitting.
The paper provides experimental results, including both quantitative and qualitative assessments, showcasing the superior performance of the framework. The results clearly highlight the effectiveness of the proposed method in facilitating personalized image generation.
Weaknesses
In Figure 3, the method utilizes Masked Language Modeling module to enhance the identity details. However, it is unclear how it would perform with fine-grained subjects (two dogs or two cats with different breed). Also, I wonder how this module would work when the concept size increases. Clarification is needed on whether the module can effectively manage such fine distinctions and multiple diverse subjects.
Recent methodologies [1, 2] have demonstrated the capability to learn multi-concept personalization, it remains uncertain if the proposed work can handle multiple personalized instances (> 2), particularly for contexts involving up to five subjects. The absence of qualitative results for three or more subjects in both the main text and appendix might be a notable omission. Including these results would substantiate the method's capability in more complex scenarios.
[1] Liu, Zhiheng, et al. "Cones 2: Customizable image synthesis with multiple subjects." arXiv preprint arXiv:2305.19327 (2023).
[2] Yeh, Chun-Hsiao, et al. "Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition." arXiv preprint arXiv:2402.15504 (2024).
Questions
Given the concerns mentioned, particularly around the method's scalability to more complex multi-subject personalizations and the clarification behind the Masked Language Modeling module, I recommend a "marginally below the acceptance threshold" for this paper. Enhancements in demonstrating multi-subject capabilities, clarity in embedding visualization, and justification for the choice of technology could potentially elevate the manuscript to meet publication standards.