Summary
In this paper, the authors give a comprehensive review of recent work on using LLMs in strategic reasoning environments (e.g., gaming environments, economic simulation). This area of research has gained much traction in the last year (e.g., all of the papers they cite in Figures 2 and 4 involves work done exclusively in 2023 and 2024) and is integral to new applications of LLMs in neighboring disciplines such as economics, business and many others. I therefore think that a survey paper of this kind is very timely and important and I generally found their paper to be easy to read and accessible.
As with any broad survey article, certain liberties were taken with how to classify different work and concepts. For example, they exclude generative agents [Park et al] and work like this as being out of scope given that the lack of focus on goal-oriented task modeling (I’d like to see more discussion of this). They also attempted to quantify the different types of reasoning skills involved with strategic reasoning tasks versus other tasks and I found this part to be overly subjective without further discussion. Given that virtually all of this work involves prompt-based LLMs, distinguishing between “prompt-engineering” approaches versus other approaches, virtually all of which involve prompting, to be a bit confusing (in fairness, the authors point this out when they write that “it is important to note that the boundaries between the above methodological categories are not entirely orthogonal”.)
Given that this article is only a survey article and offers no new empirical or technical results, I don’t have much to criticize outside of some presentation-specific issues that I note below. Baring these concerns, my main concern is about whether a survey paper fits within the COLM technical track (I don’t think it does, I will say more below. I am willing to be convinced otherwise if other reviewers and the area chair(s) are not concerned about this).
Reasons to accept
- A comprehensive and very easy to read review of LLMs and strategic reasoning, which is motivated by the considerable interest in the topic in the last year. I think it provides a useful overview of this emerging area and that researchers will cite it for this reason.
Reasons to reject
- No new technical or empirical results; not original research but a survey article. Given the call for papers (https://colmweb.org/cfp.html) and the reviewer guidelines (https://colmweb.org/ReviewGuide.html, in particular, the stated goal of having a “technical deep, exciting, forward-looking, insightful and impact program”), I don’t the paper meets the criterion for inclusion in COLM.
- (a complicated criticism) While their review is comprehensive, it fails, in my mind, to communicate what the big problems are in the different application areas where strategic reasoning has been investigated. For example, in the scenarios involving economics or game theory, what are the big or open problems in these fields that motivate using language models in place of traditional tools? What are the prospects of success in using LLMs for these problems? Of these problems, are there particular sets of problems that the LLM field has tended to focused on, or ones that people are not paying attention to?
A lot of the discussion along these lines is left vague (e.g., they end the “scenarios” section by saying: “each category [or application area of LLMs] offer[s]… unique insights and challenges”. I was really keen to get deeper into these challenges). I think it would benefit the LLM community to know a bit more about what the core problems are in these different to better motivate interest in strategic reasoning. It might also make the article easier for others outside of LLMs to engage with this work.