Summary
**Please note that this is an emergency review, completed in a limited amount of time whilst I am traveling.**
This paper studies the properties of a class of attention-based networks, modified from the standard Transformer, in the limit in which depth and width tend proportionally to infinity. Technically, it extends recent work by Li, Nica, and Roy (NeurIPS 2022) on SDE-based descriptions of MLPs with random weights. The central result is a modification of the attention mechanism, referred to by the authors as shaped attention, that gives a non-trivial limit.
Strengths
In my abbreviated reading, I found the paper clear and enjoyable to read. I think it is a worthy contribution to the rapidly-growing literature on double-scaling limits of deep and wide networks, as it takes a substantial step closer to the architectures used in practice relative to prior works. The non-commutative scaling regime of ResNets studied by the authors nicely compliments recent work by Hayou and Yang on commutative infinite-width and -depth limits of residual networks. I found the observations regarding finite-time numerical instabilities to be particularly interesting.
Due to time constraints, I have not had the chance to go through the proofs in detail.
Weaknesses
To my mind, the main weakness of this work is that it does not address inference. However, this reflects the larger issue that it is challenging to perform inference in conditionally-Gaussian processes (perhaps outside the nearly-Gaussian limit, where things can be studied perturbatively). Thus, this limitation does not significantly dampen my enthusiasm for the authors' work.
Questions
- Lines 46-47: In discussing the perturbative regime, work by Dyer and Gur Ari (2019) and by Zavatone-Veth, Canatar, Ruben, and Pehlevan (2021) should be cited alongside the work of Roberts, Yaida, and Hanin (2021).
- Lines 48-53: For completeness, it might be useful to cite work on the proportional limit from random matrix theory, e.g. by G. Akemann and colleagues.
- Line 243: "Covariane" -> "Covariance"
- Lines 322-333: In relation to my comment about inference above, I think this paragraph would be substantially improved if the authors could offer a more concrete suggestion of how their approach could be leveraged to study training and generalization.
- Lines 381-382: Why is the arXiv version of Li, Nica, and Roy cited in place of the NeurIPS 2022 publication?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
I think the authors adequately discuss the limitations of their work.