Summary
This paper looks at whether a "demand gap" exists in tasks designed for LMs, whether this gap is larger for "simpler" models, it draws on a cognitive analogy.
Reasons to accept
The paper raises an important point, that task effects may impact different models differently, and that there's a bias in how great this effect is given model size and training size. It reveals a strength of larger models (a robustness) or conversely, a weakness in the smaller models. Researchers should take this into account both for their evaluations and for deciding which models to deploy for their tasks.
I appreciated the explicit paragraph on data contamination, and I which that were more standard that it is in NLP papers these days.
The 1st, 2nd, and 5th paragraphs in the discussion were cogent and really summarized the paper and its important conclusions well.
Reasons to reject
I find it odd that the abstract assumes a competence-performance distinction in LMs. The distinction is widely but not universally accepted for humans in cognitive science, but it has been explicitly argued for decades that connectionist models (=> modern LMs) do not support such a distinction. Supporters of connectionist models of cognition saw/see this as a virtue. The consequence of this is that properties of the models relating to their architecture, eg number of parameters in this paper, and practical challenges emerging from those decisions, are fundamental to their "competence." They cannot be disentangled. While this doesn't evaluate the paper, it does bring the general narrative, as well as analogy with humans, into question.
Take a look at Schwarz (1992, Connection Science), and Allen & Seidenberg (2013, Emergentist Approaches to Language) for examples.
Buried under 3.3, there's a short paragraph disavowing direct parallels between child development and model scaling. This is of course true. Children doesn't grow more neurons as they develop, let alone by a factor of 70. However, since model size is one of the two measures of model power that this paper looks at, this severely undermines the analogy between childhood development and model power that the authors make throughout this paper. At the very least, this caveat should be brought up into the introduction, where I first noticed this problem, but a better solution would be to downplay the cognitive analogies being made here. That would also help solve the (lack of) competence-performance distinction issue raised above.
Regarding demonstrating the "signature pattern," this relies both on the theoretical assumptions about competence-performance and also practical assumptions surrounding the demand gap. Is forced choice, which is often effectively classification, really the "same" cognitive construct as the production equivalent? They're certainly very different tasks from a traditional ML perspective. If they aren't, then what is being measured isn't really the "demand gap" as stated. They're two different competencies.
Related to whether or not they're the same task further assumes that the task setups are in fact evaluating the problem that the researchers think they are. This is particularly crucial in the forced choice setting, which the authors sort of get at in footnote 2, but they don't take it far enough.
One example of this that I found particularly compelling was the following, which looked at VQA with multiple choice answers. It turns out that their model achieved most of its performance without even being exposed to the image, which means that it was simply not actually doing VQA as intended. Instead it was picking up on biases in the Q and Q construction. It took aggressive debiasing of the data set to create an evaluation where the authors were confident that the model was actually performing VQA
Wei-Lun Chao, Hexiang Hu, and Fei Sha. 2018. "Being negative but constructively: Lessons learnt from
creating better visual question answering datasets." NAACL
There was a more recent paper making the same point with regards to BLiMP. Because of the forced choice task setup and researcher-generated sentence pairs, models seem to be able to solve BLiMP without actually engaging in grammaticality. This is particularly relevant, since the authors rely on BLiMP for one of their evaluations.
Héctor Vázquez Martínez, Annika Lea Heuser, Charles Yang, and Jordan Kodner. 2023. "Evaluating Neural Language Models as Cognitive Models of Language Acquisition." GenBench
A better "demand gap" set up might, for example, compare solving math word problems with a lot of filler text and complex scenarios on one hand and more straightforward word problems on the other hand. These are more clearly the same ML task.
Given all of the above, the 2nd past paragraph of the discussion comes off as probably nonsense. It should be removed.
Typo at the bottom of page 5 "prompted asked"