Summary
This paper introduces a dataset, namely TinyStories, which was produced by ChatGPT as well as GPT-4, and which contains only simple sentences. The author argued that with this simple dataset simpler model can produce fluent and consistent stories.
Strengths
- The subject matter of the present paper looks very interesting, namely the relationships between language complexity, the size of LM, and the quality of the text generated by the LM. My impression is that there is an impossible trinity between them.
- The TinyStories dataset would be very useful for the community.
Weaknesses
In the very first place, the general research question of this paper is very unclear to me. Given the title, the paper seems to focus on exploring how small a model can be but can still produce fluent texts. But, after I started to read the paper, I found this was not the case. Both the research question and the response were problematic. I am saying this based on the following reasons:
- The research question is ill-defined. If the title is the main research question that this study intended to answer, then it is hard to know what does "coherent" actually mean in the context of this paper. This is unclear because, later on, the paper (page 1) says Tinystories was designed to capture the ability of grammar, vocabulary, facts, and reasoning. For me, none of this is about coherent.
- Another inconsistency is that by the end of the first paragraph, the paper says the question that this paper focuses on is "Is it possible to design a dataset that preserves the essence of natural language, while reducing its breadth and diversity" (which is clearly very different from what the title says). This is again not about coherence. It is also questionable why "diversity" is not a kind of "essence of natural language". Also, interestingly, on the same page, the paper says that one of the goals of TinyStories is to train a model that can produce *diverse* stories.
- The design of the whole study looks very unscientific to me. Regardless of what exactly the research question is, this study has too many free variables, at least including, the complexity of the language in the dataset, the size of the dataset and the size of the model. In this sense, I don't think the conclusions are reliable and generalisable.
Apparently, a clear trade-off exists in the dataset, namely the complexity of the language. (btw. for me, this is a type of "essence of natural language".) I don't think it is a limitation of the dataset but do think it is an important aspect to research and discuss. Nonetheless, unfortunately, this paper seems to totally overlook this important aspect.
Given the task defined, this paper also proposed an evaluation framework, the reliability and rationality of which are both questioned: (1) it is quite unreasonable to train a model using contents produced by GPT and, meanwhile, also evaluate the same model using the same GPT. (2) only 50 prompts are used in the evaluation. Since only automatic evaluation was used in this study, I wonder why more prompts were not considered. (3) it is also unclear why only assessing grammar, creativity and consistency. None of them is about either coherence or fluency. More importantly, creativity seems to be irrelevant to any goals mentioned in the introduction.
Generally speaking, for me, the paper looks more like a position paper rather than a regular paper for ICLR as while no serious evaluation had been conducted, meanwhile, the last 4 pages are all about example outputs.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.