An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem

Deep learning models can exhibit what appears to be a sudden ability to solve a new problem as training time, training data, or model size increases, a phenomenon known as emergence. In this paper, we present a framework where each new ability (a skill) is represented as a basis function. We solve a simple multi-linear model in this skill-basis, finding analytic expressions for the emergence of new skills, as well as for scaling laws of the loss with training time, data size, model size, and optimal compute. We compare our detailed calculations to direct simulations of a two-layer neural network trained on multitask sparse parity, where the tasks in the dataset are distributed according to a power-law. Our simple model captures, using a single fit parameter, the sigmoidal emergence of multiple new skills as training time, data size or model size increases in the neural network.

Paper

Similar papers

Peer review

Reviewer TV9v7/10 · confidence 3/52024-07-11

Summary

The paper presents a model capable of predicting the appearance of emergent abilities, using only information on the emergence of the first ability. The model successfully predicts emergence in a 2-layer MLP solving the multitask sparse parity problem, a toy-model problem constructed with a power-law skill distribution, to mimic the theorized distribution of language datasets. The model also exhibits scaling laws known to exist in this problem.

Strengths

The paper expands on recent works, in a significant and active field of research. The authors present convincing results when evaluating the proposed emergence model against MLP training data. Extensive derivations are provided in the appendix, although I did not carefully check them. The paper is written in a clear manner and results are visualized in clear formats.

Weaknesses

- The scope of the experiments is too small in my opinion, specifically Figure 1. It would be nice to see results for a lot more than 5 skills, since it seems possible that predicting the emergence of less-frequent skills becomes harder at higher orders of magnitude. My impression is that the experimental setup is light enough that scaling up will not require unreasonable resources. - The general scope of the paper is a bit small, focusing only on the multitask sparse parity problem. It would have been nice to see results on more natural settings. That said, the current results on the chosen problem setting are still interesting on their own. - Regarding the scaling laws presented in Table 2, glossing over the paper gives the impression that these are new scaling laws found by the authors, when in fact they are reproductions of the scaling laws calculated by Michaud et. al. It might help to clarify the relation to Michaud et. al. in the main section of the paper.

Questions

The main section of the paper does not explain why the input data is distributed as a power law (Zipf's law), making it hard to understand the justification for readers unfamiliar with previous literature, i.e. Michaud et. al. I would suggest adding a short explanation when introducing the problem.

Rating

7

Confidence

3

Soundness

4

Presentation

4

Contribution

3

Limitations

Limitations are adequately discussed, including scope limitations.

Reviewer D3CL7/10 · confidence 5/52024-07-13

Summary

I’ll write two summaries: one to state my “moral” understanding of the work and another to state the specific contributions **“Moral” Summary:** The authors propose an analytically solvable model to study scaling laws and emergent abilities by combining the problem of Barak et al 2022 + the data distributional assumptions of Michaud et al. 2023/2024 + the model of Saxe et al. 2024 + their own innovations. **Specific Summary:** - The authors study the multi-task sparse parity problem studied by Barak & Michaud - The authors propose a multilinear model that identifies each sparse parity problem as a “skill” and then define the learned function as a “multilinear” (i.e. two independent linear parameters) function of the “skill” basis functions - In this model, one can then exactly compute scaling and emergent abilities as a function of the typical scaling parameters (compute, data, parameters) - The authors also train 2 layer MLPs and transformers to test how closely their maths match empirical results

Strengths

- Overall, I think this is a really well done paper (although I’m concerned about Figure 1 - see weaknesses). What it does, it does thoroughly. - Table 1 is useful in helping explain the multi-task parity problem - Table 2 is a great way to summarize and organize both the results as well as the conditions of each result

Weaknesses

- Emergence is a phenomenon studied at scale, and while I strongly support trying to find simplified models of large-scale phenomena, I feel that there needs to be an attempt to connect back to the original phenomena of interest. For instance, Michaud et al. 2023/2024 (Citation 17 in this paper) - on which this work most closely connects - at least attempts to look into real language model pretraining data. In contrast, this paper makes no such attempt, which I think strongly limits it. This is the #1 reason I don't feel comfortable giving a higher score. - Following the above point, I consequently think a more focused title (e.g., “An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem”) would better represent this paper. - For an exactly solvable model, Figure 1 shows that the predictions only roughly match the experimental results. This becomes especially pronounced for higher k. For instance, look at Figure 1(c) orange. The prediction is that there should be a rapid leap from 0 to 1, but instead, there is a sigmoid-like transition with long tapering tails. Green, red and purple are all similar. The inability to predict 2-layer MLPs makes me think that this analytically solvable model is already only an approximation of incredibly simple networks

Questions

Figure 4: What is the timescale of each skill’s emergence for the transformer? I can’t tell if the higher k lines are step functions or sigmoidal functions that are compressed by the log scaling of the x axis. To be clear, I’m not asking for when the skills emerge, but how long it takes for each skill to emerge.

Rating

7

Confidence

5

Soundness

4

Presentation

4

Contribution

3

Limitations

N/A

Reviewer owz96/10 · confidence 3/52024-07-15

Summary

This paper provides an in-depth study of a toy dataset and model, both in terms of scaling laws and emergent skills. They study the ‘multitask sparse parity’ synthetic problem introduced by Michaud et al. (target is the parity function of a string of random bits, each task indicates a different subset of bit ids, task frequency follows a power law). New in this work, they consider a ‘multilinear’ model (i.e. $y=ab x$ rather than simply $y=cx$) to incorporate the dynamics of a two layer neural net (as per Saxe et al. in a different setting). The paper recovers scaling laws (which agree with Michaud et al.’s coefficients for T, D & N, and additionally adds for C=TN). The model also gives rise to emergence of each skill with a sigmoidal shape. Rarer skills are learned later on.

Strengths

- From a technical perspective, the paper is a very strong and complete piece of work. It exhaustively sweeps through results and surrounding analysis of the setup studied in the paper (even, for example, having proofs of the scaling laws in increasing resolution). - I have confidence in the claims and analysis. I spot checked several parts in depth, however, I have not been able to interrogate all claims and analysis in the paper (the appendix pushes the paper to 54 pages). - Providing a model that unites sigmoidal-shaped skill emergence with scaling laws, is a tantalizing prospect, as these are two major properties of LLMs. A theoretical model combining the two aspects would be of high interest to several parts of the community.

Weaknesses

- The paper contains a huge amount of material. A lot of the good stuff is buried in the appendix (related work is important, the contrast with a linear $y=cx$ model, the scaling law proof sketches, the stage-like training discussion). The reading experience suffers from this, and a conference paper struggles to do it justice. It's possible it would be better suited to a long-form journal format (and would allow reviewers more time to comb through the details). - The scaling laws derived by the paper match the coefficients found in Michaud et al. (with the addition of $C$). There is a slight difference in that the setup now uses $y=abx$, but I generally felt that given the repeated outcome, providing scaling laws derived with varying resolutions need not be such a major focus of the paper and appendix (maybe some nuance has been lost on me?). Reducing this would free up bandwidth to allow readers to absorb the other more interesting parts. - A concern I have is in how realistic the setup of the dataset and model is. The emergence and scaling laws come from a lightweight two-parameter linear regressor which receives ideal task indicators as input. The closest interpretation in a real setting I can think of, is as a two-layer linear MLP head, placed on top of a deep pretrained model that is frozen with powerful representations for each task already learned. This paper is still valuable, but I might suggest a more scoped title because of this – ‘an exactly solvable model for emergence and scaling laws’ is not inaccurate, but might be a little broad.

Questions

See weaknesses.

Rating

6

Confidence

3

Soundness

4

Presentation

3

Contribution

2

Limitations

Fine.

Reviewer NXT53/10 · confidence 3/52024-07-16

Summary

The paper proposes to use a certain generalization of the well-known sparse parity problem, as a theoretical framework for neural scaling laws. The theoretical setup doesn't seem to make sense (I'm open to changing my mind). Indeed, equations (2) and (4) taken together give \begin{equation*} f^*(i,x) = S\sum_{k=1}^{n_s} g_k(i,x) = S\sum_{k=1} \delta_{ki}g_i(i,x) = S g_i(i,x). \end{equation*} This is definitely not "a sum of $n_s$ skills" (as claimed by the authors); it is a single skill. The same issue repeats itself in (9) when the authors introduced their so-called multi-linear model. Indeed, that model simplifies to (again thanks to (2)) \begin{equation*} f_T(i,x) = \sum_{k=1}^{n_s} a_k(T)b_k(T)g_k(i,x) = \ldots = a_i(T)b_i(T)g_i(i,x). \end{equation*} From this point onward, it is not clear what the paper is trying to do. Second, it is not clear what the paper is trying to achieve beyond what was already done in Michaud et al's "Quantization Hypothesis" paper (for the record, the paper proposed the multi-task sparse parity problem as a example exhibiting scaling laws w.r.t sample size and model size, within the framework of their quantization hypothesis). Finally, the paper seems to be missing some important literature, for example Cabannes et al. "Scaling Laws for Associative Memories" (ICLR 2024), which proposes finite-capacity extension of Hutter's model, and establishes an array of different scaling laws for different learning algorithms.

Strengths

As explained already, my low score is because I think the theoretical setup of the paper is unclear. I'm open to changing my mind.

Weaknesses

As explained already, my low score is because I think the theoretical setup of the paper is unclear. I'm open to changing my mind.

Questions

In what way do (3) and (9) represent "a sum of skills" ?

Rating

3

Confidence

3

Soundness

1

Presentation

2

Contribution

1

Limitations

Yes (in Section 5.4)

Reviewer NXT52024-08-13

- Thanks for the clarification; a notation problem indeed. - I have a good understanding of the paper now. I think the contribution is interesting but still incremental (based on current literature on provable scaling laws). I'm increasing my score to 6.

Authorsrebuttal2024-08-13

Dear reviewer, We are glad that we clarified the confusion and thank you for updating the score.

Area Chair bVPQ2024-08-13

Dear reviewer, Thank you for your review and efforts. Please update your score on the original review as well.

Reviewer D3CL2024-08-08

Brief Comment on Contributions beyond Michaud et al. 2023

I'll respond to the authors tomorrow, but briefly, on this overall comment, I'd like to clarify one point concerning "Contributions beyond Michaud et al. 2023" today in case other reviewers look at this in the interim. I personally think that this work extends far beyond the Michaud et al. 2023 paper. If any other reviewers feel differently, please explain why or please point me towards where you might have already explained in your review. Thank you!

Reviewer TV9v2024-08-11

I thank the authors for their detailed response. In light of the changes proposed, I will change my score to accept (7). One question about the experimental setup: What were the hardware requirements for the 2-layer MLP experiments in figure 1?

Authorsrebuttal2024-08-12

Dear reviewer, we deeply appreciate updating the score. The specification of setup is detailed in Appendix K.5. Each run of the experiment – one point in the figure – requires 2 to 5 hours for time emergence and 20 to 50 hours for other experiments on a CPU. We ran the experiments on a CPU cluster because the small size of the MLP allows less pronounced difference in the running on GPU and CPU (running on RTX 4090 GPU was typically only 3X faster than an average CPU in the cluster) while we benefited from parallelism in a larger CPU cluster. The most demanding experiment was Fig.1(b) with data emergence. The figure has 30 different numbers of datapoints ($30$ runs), repeated $10$ times for the error bars. We ran for $5 \times 10^5$ steps for data and parameter emergence (compared to $3 \times 10^4$ for time emergence) to remove the potential effect from early stopping (i.e. to assure $T \gg D$). To observe the emergence of an additional skill, we require a magnitude increase in $D$ which leads to a magnitude increase in $T$ – to assert $T \gg D$ – and a magnitude increase in the batch size – to mitigate the SGD noise.

Reviewer D3CL2024-08-12

Response to Authors' Rebuttal

Thank you for your comments! I appreciate the additional figures in the global response and your answers to my questions. I'm going to increase my confidence but keep the score. I don't feel comfortable increasing my score because while I feel this paper is very thorough, I feel it is limited in its general applicability without strong connections back to more realistic (larger, more data, real data) models.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC