Summary
The authors of this paper examine the performance improvements of language models over the past decade, and investigate how much of it can be attributed to algorithmic improvements of language models.
Strengths
I note here that, given that this paper proposes a method to evaluate the historical progress of language models, I believe that the primary goal for the paper would be to provide interesting insights for the field and outline current open questions. As such, I will structure the rest of my review with that in mind.
- The topic of language models is more central than ever to the broad NeurIPS community, and I think that analysis on the historical progress of the field is of great interest.
- The authors obtain valuable insight on the cause of language model improvements over the years. I find their conclusion that data/compute scaling has been a major driving force for improvement over the past few years interesting, if somewhat expected.
- The result on the importance of the transformer is also interesting, and provides a nice retrospective justification of the broad adoption of this architecture.
- The analysis of data from past works is also sound.
Weaknesses
- The most crucial weakness of this paper is that, despite the analysis of past trends, it does not provide clear insights or suggestions on where the field should direct its efforts, moving on. The authors acknowledge this limitation, but it seems to me that such discussion Is crucial, in order for this paper to be of interest to the community.
- It is also not clear to me how the insights provided by the paper could extrapolate in the future. While there has been a lot of effort and improvements gained in performance from data scaling, this is not something that can reliably keep on - at some point, sources of data and compute limitations catch up. As such, it is not clear to me how the insight that the effective compute for language models doubles every set period of time can be useful for many years down the line.
- On more minor note, I think some points can be improved for clarity of presentation:
- The captions of Figures 1a and 1b are joined together - some spacing between them is required.
- It would be better if the authors clarified the statistical models used in the main paper, rather than in the appendix.
Overall, my key concern for the paper is that it does not provide enough insight for future directions.
Questions
I would be grateful if the authors could expand on the points I raised above, regarding the insight for future directions.
Limitations
The authors have adequately addressed the limitations of their work. Regarding negative societal impact, I do not foresee any arising from this work.