Summary
The presented approach gets much better results on key measures of creativity compared to SOTA.
This paper is a well-written presentation of the novel LLM Discussion approach to generating creative responses on established creativity and scientific innovation benchmark tests.
The approach is a highly original simulation of "collaborative discussions with diversified peers" that involves (i) a three-phase, initiation-discussion-convergence framework for multi-turn dialogues designed to elicit diverse and creative responses from (ii) LLM-based agents, each of what are assigned distinct roles and are directed to generate answers building on the others' responses.
This framework, (i), is compared against single-LLM and multi-LLM approaches and outperforms these previous methods on several creativity metrics.
The role-playing LLM-based agents, (ii), cleverly leverages the ability of LLMs to imitate very different personalities in order to generate more diversity of ideas.
Reasons to accept
* Ambitious objective to simulate "collaborative discussions with diversified peers" with novel design of 3-phase framework for discussions by role-playing LLM agents.
* Exciting results with strong gains over previous methods of generating creative answers.
* Clear delineation of where this approach contrasts with state of the art.
* Solid empirical evaluation methodology with established tests and metrics, engaging both automated and human judgements, as well as documented ablation tests, explaining development of specialized prompts, # discussion rounds and # LLM agents
Reasons to reject
* The authors do not discuss going beyond the role-playing agents as artificial people, and address other ways of coming up with better answers,.
* The paper does not address what is actually happening inside the model, that would also give deeper insights.
Questions to authors
* The space of possible prompts and ways of combining the results of prompts is so broad that it seems unlikely these results are anywhere near the best that could be obtained with present models. Q1: Why, for instance, are the particular roles of "visionary millionaire, startup founder, social entrepreneur, creative professional, customer, environmentalist, digital nomad, industry insider, futurist" (which sounds like a list of people you might encounter at a party in Silicon Valley) chosen instead of, say, homemaker, mathematician, janitor, and professional actor? Or, say, all 16 Meyers-Briggs personality types? Or people from 20 diverse countries or time periods? There's no justification for this and no way of knowing what would be best a priori. At the moment there is so much low-hanging fruit that it is easy to come up with new approaches that get big gains, but it would be helpful to find some automated way of searching the space of possibilities more thoroughly.
* Q2: What might help explain why fluency and flexibility scores are lower for LLM Discussion than LLM Debate, even though these matter less than originality and elaboration as measures of creativity. Perhaps there would be some simple way to alter the prompts to boost these factors as well?
* Q3: how is the prompt in the convergence phase leading to "convergence" as opposed to simply enumerating answers from prior phase, given instruction to "present a list"? Did you experiment with an explicit prompt to create a summary or draw a conclusion?