Summary
This paper presents various separation results, showing that a Transformer with CoT can solve certain formally-defined reasoning tasks, but a Transformer *without* CoT cannot (assuming bounded depth). This sheds light on the power of CoT. The formal results are supplemented with empirical results that support the claims.
Strengths
- Understanding CoT is an important and timely question. This paper is therefore tackling a significant question.
- To the best of my knowledge, this is the first paper to provide a theoretical explanation of the power of CoT (i.e., originality).
- In general, the paper is quite clear and high-quality, though I think the notation could be improved (see below).
Weaknesses
- I think calling it CoT but focusing on "CoT generation" is perhaps misleading. Many people think of CoT as a prompting technique. In my opinion, the generation aspect that this paper focuses on is actually bigger/more important than CoT suggests, and I actually think that - although buzzwordy - CoT diminishes the general power of generation that this paper is getting at (i.e., this paper transcends CoT). I would suggest looking into alternative, more general titles.
- Section 4.1 has a lot of notation, including double subscripts. I think there's a significant lack of accessibility created by the heavy notation. I would really encourage the authors to think about whether the notation could be simplified. I think doing so could substantially increase the long-term impact of this paper.
- Relatedly, equations (4) and (5) would be clearer if their new objects were introduced and explained conceptually before jumping into equations (4) and (5).
- Overall, I think the paper could do a better job explaining the significance of the theoretical results. Although this is an important problem to study theoretically, and this paper does a good job of initiating that study, it's not entirely clear what can be gained conceptually from this analysis. I think many people already intuitively grasp that generating more tokens gives an LLM more power, and various methods that allow more tokens to be generated before arriving at a final answer are more powerful. It would be nice to understand whether there's a conceptual message here that goes any deeper than the aforementioned intuition.
Questions
- See CoT naming comment above. Do you agree?
- See notation comment above. Is there a way to simplify the notation in Section 4.1? Is there a reason it has to be this complicated? If it seems necessary, maybe there's a simpler version that can be presented in the body of the paper, with the full version moved to the appendix?
- My main question is also discussed in the Weaknesses section. Is there any conceptual takeaway from this theoretical analysis beyond "generating more tokens is more powerful"? Even without that, I think this is a strong paper. However, I think this could be an especially powerful paper if the intuitions from the theoretical analysis could be further synthesized in this direction (and made accessible to readers).
Rating
8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
The authors do a nice job of discussing some of the limitations of this work. There is no discussion of societal impact, but I do not think this is a problem for this particular paper.