Acknowledging my wrong interpretation on some points, but confirming other doubts.
Thanks to your comments I was able to dissect better the paper, and I will increase my score to **Reject** (and also reduce my confidence on this paper to 2 from 3).
With self-doubt, I have re-read the paper from scratch again, going over the appendix.
As a result, I understand that your frameworks evaluate whether a collusion might appear in practice, by looking at the exchange of messages between two LLMs. For this reason I will rise the score as stated, and I apologize for not having understood the paper in the first run.
However, I still think that this paper has issues that prevent acceptance as follows.
**The paper is not well-written, since all major content is in Appendix** In my second run, I still had to re-read the paper multiple times, and I got **what** has been presented from the Appendix, and not the main paper. For instance, all the single and group tests were unclear on their setup, goal, and execution. Thus, authors should better explicit both steps in main paper, and only later clarify minimal details.
I would reduce section 4 (maybe included in sect 3 as well) and I would expand section 5 to clarify that:
* you instruct LLMs with system prompts to assess whether they can use steganographic techniques, like base64 or chiper algorithms (this is only hinted in one sentence at page 6).
* how you apply group evaluations, since in Appendix H is only written that models are applying techniques based on single algorithm generation (point 1), but I still don't understand how it works. Again, the explanation on the precise setup must go in the main paper.
**Benchmark of steganographic techniques.** While it is interesting that LLM can generate code for creating steganographic content, I do not think that such stands on a contribution on its own. In fact, all the presented techniques are wide known in public literature, and they are generated by LLMs just as code for many languages. This is similar to previous work showing that LLM can write code. An higher contribution would have been showing that previously it was not possible, or that alignment and guardrails prevent the generation of such content.
**I still think experiments are misleading.** By reviewing the paper, it is hard to understand the message that these figures are conveying. I can understand what is happening only after having read all the Appendix. If I got it correct, the success rate is the relative number of times that an LLM generated the required content (Fig5). Thus, it appears that only GPT4 (which is closed source LLM) is capable of generating all the techniques, while all the other baseline fail to even reach 30% (which are 3 attempt over 10 on average). From my point of view, this does not look as a contribution, since the only capable one is a closed source LLM that might rely on extra knowledge, and there is no way for telling the contrary.
**Overseer?** In all the above, I miss the relevance of the overseer presented in the figure, since I can't understand the setting of group evaluations.
I truly apologize with the authors for not having understood parts of the paper in the first run, but re-reading again confirmed some of my belief about contributions, setting up of experiments, presentation of content.