Reviewer Lu7o Replies
<1. baselines>
- From your rebuttal, I have understood what is your key advantage you want to show now. Precisely, that is the **efficiency** compared to other baselines (w/ your similar level of performance as many baselines), such as AutoPrompt, RLPrompt, etc. It would be much better for you to consider adding efficiency comparisons visualizations/discussions. Additionally, I am not sure whether RLPrompt would fail in this case, as image-level/instance-level optimization takes fewer time, which deserves some experiments as well. If it is weaker/much slower than your approach in a few cases, it would be good to include this and highlight several times in your paper, to make it more clear about your key advantages. Moreover, from your new AutoPrompt experiments, it would be good for you to add one **rigorous** figure on the efficiency comparison results, e.g., using the best setup and also the less good setup (such as other LR, prompt length in AutoPrompt), to position your approach. This would definitely add some values to your paper.
<2. CLIP interrogator>
- Thanks for your clarification. It does seem that interrogator performs much weaker with fewer tokens. But as you said, it would be good for you to incorporate some efficiency studies as well. That can address my concerns on your approach vs. CLIP interrogator.
<3. Secret Language>
- This part is just for one of your claims. You said you identify a secret language of text-to-image diffusion models [1]. But in the initial paper, the lexical items that they used are purely gibberish tokens, w/o any significant real-world meanings, such as "Apoploe vesrreaitais" for birds. But in your results, and also your new pdf results, you always include some natural tokens, for instance, "butterfly emoji + users', 'cruise + green render lights', "balloon relationships", “cat” for cat image, and many more shown in your paper figures. I am curious about more fine-grained studies on your generated prompts. For instance, others might suspect that when deleting your natural meaningful tokens, your generated results would fail completely, which may indicate a weaker claim of "secret language". Or in other words, perhaps, most of your used nonsense tokens are mainly search artifacts. Therefore, my comments are to say it would be better for you to provide some rigorous studies.
I believe this point is also mentioned by reviewer DXf6, for which I agree with him/her.
[1] Discovering the Hidden Vocabulary of DALLE-2
<4. efficiency results>
- See 1, it would be better for you to provide rigorous studies to position your approach.
For current rating, I am saying that this work should be further improved to be an interesting paper published in this conference. So I select the borderline rating, in which I would like to wait for other reviewers or ACs for final judgements.