Double and Single Descent in Causal Inference with an Application to High-Dimensional Synthetic Control
Motivated by a recent literature on the double-descent phenomenon in machine learning, we consider highly over-parameterized models in causal inference, including synthetic control with many control units. In such models, there may be so many free parameters that the model fits the training data perfectly. We first investigate high-dimensional linear regression for imputing wage data and estimating average treatment effects, where we find that models with many more covariates than sample size can outperform simple ones. We then document the performance of high-dimensional synthetic control estimators with many control units. We find that adding control units can help improve imputation performance even beyond the point where the pre-treatment fit is perfect. We provide a unified theoretical perspective on the performance of these high-dimensional models. Specifically, we show that more complex models can be interpreted as model-averaging estimators over simpler ones, which we link to an improvement in average performance. This perspective yields concrete insights into the use of synthetic control when control units are many relative to the number of pre-treatment periods.
Paper
Similar papers
Peer review
Summary
This paper investigates over-parametrized models in causal inference, specifically focusing on high-dimensional linear regression models and high-dimensional synthetic control estimators with a large number of control units. The authors examine the prediction risk behaviors associated with these two estimators, and present a unified theoretical perspective on the descent phenomena in the interpolating (=high-dimensional) regime, which they refer to as the “model-averaging property” as demonstrated in Proposition 1 and Proposition 4; also see Eq. (MA). By introducing the high-level assumption of the model-averaging property (MA), which posits that the full model can be expressed as a convex combination of simpler “leave-one-out” counterparts, and another technical assumption (P) concerning the optimality of the aggregated model against a random permutation of the weights, the authors assert that the complex model cannot perform worse than the average performance of the simpler models (Proposition 5). In summary, this paper provides a concise exposition of the “benign overfitting” phenomena observed in two estimators widely utilized in causal inference. The authors rely on two high-level geometric assumptions, the model-averaging property and Assumption (P), which provide insights into the descent behavior of these estimators and warrant further investigations in follow-up studies.
Strengths
This paper offers a simple yet intriguing perspective on understanding the double descent phenomenon in the interpolating regime by leveraging the mechanical properties of predictive estimators that are largely agnostic about the underlying data generating process. The authors adeptly combine knowledge from linear algebra and convex analysis to provide concise explanations for the returns of complex models in high-dimensional causal estimators. The practical implications of their findings, particularly in the context of synthetic control with many control units, are noteworthy. The paper alleviates concerns about overfitting when using a large number of control units, eliminating the need for pre-selecting an appropriate subset, which could be challenging in practice. Furthermore, the authors' focused approach on linear regression models and synthetic control estimators enhances the clarity and concreteness of their arguments. They effectively develop their findings in these two settings first, and thereafter, hinting at the potential extension of the abstract results based on the model-averaging property (MA) and the permutation property (P) to broader settings beyond the two examined scenarios. While it is yet challenging to assess the significance of this work within the broader NeurIPS readership, I believe the authors make a meaningful contribution to the sub-community of econometrics/causal inference. The paper offers a fresh and clear perspective on understanding the phenomenon of benign overfitting, delivering an important message that reassures the practical utilization of high-dimensional synthetic control, specifically.
Weaknesses
Although this paper has several strengths, there are three concerns/suggestions that could be addressed to further improve its quality. Firstly, there is a concern about the possibly limited scope of applicability for the presented results. Although the paper establishes theoretical support for the concept of "benign overfitting" in synthetic control estimators and acknowledges the possibility of extending the analysis to a broader class of estimators through the abstract conditions (MA) and (P), it remains unclear how far-reaching this perspective can be and to what extent it can be extended. It would be valuable if the authors could provide a discussion on the expected scope of extensions, limitations, and potential challenges that may arise. Additionally, in order to strengthen the positioning of the paper, it would be beneficial to provide additional comparisons and contrasts to previous approaches. Given the focus on synthetic control estimators, the authors could comment on the limitations of applying existing results on double descent and benign overfitting in over-parametrized models to the specific context of synthetic control estimators. This would effectively highlight the authors' contributions in this work and emphasize the distinctive features of the techniques employed. Lastly, while the paper primarily presents a novel theoretical perspective, it would be valuable to complement the theory with a more extensive set of numerical experiments. For instance, conducting an ablation study with synthetic datasets to verify the assumptions and theory, as well as performing experiments with real-world datasets at scale to confirm the descent behavior of the synthetic control method, could strengthen the paper's contributions. By incorporating such experiments, the authors can provide empirical evidence to support their theoretical findings and enhance the practical relevance of the paper. By addressing these concerns, the authors would be able to further strengthen their contributions in this paper, clarify the scope of their results, highlight its distinctive features, and provide a more comprehensive explanation for the implications of their findings.
Questions
1. I would like to request the authors to address the three concerns raised in the “Weaknesses” section. 2. Miscellaneous suggestions: (a) It would be clearer if the authors replace the term "full rank" with "full row rank" for clarity, e.g., in lines 127, 171, 174, 191, although it is obvious in the interpolating regime. (b) In lines 153-155, the authors state, "In this case, loss continues to decrease throughout, ultimately reaching a minimum that is below the lowest error achieved left of the interpolation threshold." However, this is not easily discernible in Figure 1-(b).
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
2 fair
Limitations
The paper suggests several potential avenues for future research, recognizing some limitations of their findings, such as exploring more fundamental conditions beyond (P). However, it could benefit from a more explicit discussion on the limitations and potential negative societal impact of the proposed approaches. It would be valuable for the authors to provide a more detailed acknowledgment of the technical limitations of their methods and offer insights into the potential adverse consequences that may arise when applying these approaches in real-world settings, specifically in the context of synthetic control methods. Nonetheless, it is important to note that given the paper's primary focus on theoretical aspects, a comprehensive and extensive examination of these limitations may not be deemed critical.
Summary
The paper studies the issue of double and single descent for linear regression as well as synthetic control. For the former, empirical motivation is given in the context of predicting wages and theoretical results are given from the perspective of model averaging. For the latter, empirical motivation is given in the context of imputing counterfactual California smoking rates and again theoretical results are given from the model averaging perspective.
Strengths
- To my reading, the most impressive result is Proposition 4, which shows that synthetic control has the model-averaging property. I presume that this is new in the literature but it is not explicitly mentioned in the paper. It would be good to clarify the novelty of Proposition 4 and emphasize its implications. - The two numerical examples are illustrative and provide very good empirical motivation.
Weaknesses
- To my reading of the literature, one of the most under-explored but central issues in double descent or benign overfitting is the analysis of bias. For example, Tsigler and Bartlett (2023, JMLR) entitled "Benign overfitting in ridge regression" provides sharp bounds for the bias term. The current paper does not provide any result regarding the bias term, which will not be zero with over-parametrization. - It seems that it is not totally new to combine model averaging with double descent in the literature. For example, see the following quote from Wilson and Izmailov (2020, NeurIPS, page 8) entitled "Bayesian deep learning and a probabilistic perspective of generalization": _"Double descent [e.g., 3] describes generalization error that decreases, increases, and then again decreases, with increases in model flexibility. ... However, our perspective of generalization suggests that performance should monotonically improve as we increase model flexibility when we use Bayesian model averaging with a reasonable prior."_ It would be helpful to provide a more thorough discussion of the literature.
Questions
- Tsigler and Bartlett (2023, JMLR) gives general sufficient conditions under which the optimal regularization parameter is negative. In view of this, I am wondering what would happen if $\eta$ in equation (2) is negative. In other words, would it be possible to study the penalized synthetic control estimator with a negative regularization parameter? - Model averaging is popular in both statistics and econometrics: e.g., Claeskens and Hjort (2008), Model selection and model averaging, Cambridge University Press; Hansen (2007), Least Squares Model Averaging. Econometrica. Some discussion of related papers on model averaging would be helpful.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
2 fair
Limitations
- Lines 221-223: it is stated that "We believe that these variance and geometric properties of linear regression are well understood in the literature and likely not new, although we are not aware of an explicit statement of the model-averaging connection between more and less complex interpolating linear-regression models." The statement is a bit unclear. It would be better if the literature review is more thoroughly done in the paper.
Summary
This paper examines single and double descent phenomena in two causal inference estimators: high-dimensional linear regression and synthetic control estimators with many controls. To begin, the paper starts with a high-dimensional linear regression problem and illustrates the double descent phenomena using the famous LaLonde dataset. Then, they show that complex models can be seen as the model averaging of simple models when using interpolating linear least squares estimators. Their main contribution is to explore the single and double descent phenomena in synthetic control methods for the first time in the literature. While it is commonly recommended not to include too many control units, the authors found the single descent phenomena: the performance monotonically increased as the number of control units increased. They derive the model-averaging-based risk bounds to explain the single phenomena they observed in the real-world example.
Strengths
**Originality** To the best of my knowledge, this is the first paper to explore the single and double descent phenomena in causal inference settings. In particular, for synthetic control problems, there are a large number of control units in many important settings. For example, when a company tries to estimate the effect of a certain internal policy on its performance, there are potentially a large number of firms they can use as the donor pool. The conventional recommendation is to focus only on a small number of selected control units, but this paper opens up a new opportunity for researchers to incorporate a large number of control units beyond the number of pre-treatment periods allowed. **Quality** This paper provides a very useful insight into the single and double descent phenomena from a model averaging perspective. Providing both intuitive and geometric interpretations of results was helpful. **Clarity** This paper is clearly written, and starting with a simpler linear regression as a basis for the synthetic control method was also a very effective presentation. **Significance** As mentioned above, it will be significant as this paper will open up a new opportunity for researchers to incorporate a large number of control units beyond the number of pre-treatment periods allowed.
Weaknesses
I do not think there is any obvious weakness in the paper with respect to the goal of the paper. I have some suggestions about making the paper more relevant to realistic causal inference settings, and I list them as clarifying questions below.
Questions
**(1) Comparison to a "reasonable" simple model in a bit more realistic setting** In many deep-learning settings, it is probably difficult to think about a "reasonable" set of covariates as the variables might have very little substantive meaning in some applications like image detection. But, for many causal inference problems like those examined in this paper, there are many "simple" models that any social scientist would start with. While this paper provides very interesting insight about the single and double descent phenomenon, I think the paper would be even stronger if the authors can demonstrate the idea of double descent and over-parametrization can actually beat simple yet reasonable models. As far as I understand, in both empirical examples (Figure 1 and Figure 4), the baseline model is to include variables "randomly" (random simple models) and then average the performance across such random simple models. And the average of such random simple models can have low performance as it averages over many "random" non-reasonable models. I want to emphasize that the original authors already acknowledged this point in the paper, and I am here to encourage the authors to explicitly include some empirical evidence about this point. (1.1) LaLonde data In the LaLonde data example, because many variables are constructed just from eight variables, it is expected that most of the 8000 variables have very low signals and a random subset of such variables can be far from informative variables. **Can the authors include the RMSE based on a simple linear regression that includes the original eight variables additively?** Can overparameterization outperform this baseline model? (1.2) Synthetic Control Problem It is interesting to see that it only shows the single descent phenomenon. But I was wondering whether this is due to the very small number of pre-treatment periods (3). Essentially the regime before the interpolation threshold could be too small to show any bias-variance tradeoff in a meaningful sense. **Can the authors include at least ten pre-treatment periods, which is often recommended in practice?** Do we still see the single phenomenon with a larger number of pre-treatment periods? **(2) Overparametrization vs Regularization** Another simple benchmark is a classical regularized model. For example, for the Figure 1 problem, **can the over-parameterized model beat a simple Lasso?** For the linear regression problem, the minimal-norm interpolation solution picks a solution that has the minimum norm rather than solving the objective function with a penalty term. **What is the theoretical connection between a classical penalized regression (adding a penalty to the objective function to make a solution unique) and the over-parametrized model (finding all the solutions that make the objective function equal to zero and finding one solution that has the minimum norm)?** **(3) Inference** This paper considers the RMSE of causal estimation. But, in many causal inference settings, researchers are also interested in estimating confidence intervals. I am curious to know whether **over-parametrized models** will make inference intractable or can be addressed similarly in the double machine learning or the semiparametric inference literature (for a linear regression case; as far as I know, there is no unified inference framework for the synthetic control method yet). **(4) Connection to Super Learners** SuperLearners by van der Laan and his colleagues also use the convex combination of individual prediction methods to improve the performance of the ML prediction. **Is there any connection to SuperLeaners?**
Rating
9: Very Strong Accept: Technically flawless paper with groundbreaking impact on at least one area of AI/ML and excellent impact on multiple areas of AI/ML, with flawless evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
4 excellent
Presentation
4 excellent
Contribution
3 good
Limitations
They clarify the limitations of their approach.
Summary
The paper studies overparameterized linear regression in the context of imputing data for estimation of causal effects. This involves both unconstrained regression for imputing wages in the CPS dataset and regression under a simplex constraint on the weights for synthetic controls in the California smoking rates dataset. Favorable performance is observed for the overparameterized estimators, where the number of donors is larger than the number of samples. The main contribution of the paper is an explanation of this favorable performance. It is shown that minimum norm solutions in overparamterized linear regression can be expressed as weighted averages of least squares solutions under subsets of the donors. Then it is claimed that such averaging results in better generalization in imputing the missing values.
Strengths
I think that problems regarding the use of overparameterized models in causal inference are very important and interesting. The paper makes simple yet non-trivial observations about the model-averaging property of min-norm linear regression models when used in causal inference tasks. It offers some observations that might prove useful in practical considerations on whether one should include as many donors as possible versus selecting donors carefully from a large pool. Finally, I found the presentation quite simple and fluent. Overall, reading the paper was an enjoyable experience.
Weaknesses
The main weaknesses of the paper in my humble opinion are as follows: 1) I think that the explanation for why model averaging results in better generalization is somewhat lacking. The closest result to a generalization bound is proposition 5, though this only proves the resulting classifier is better than the worst-case classifier. It is quite far from results on benign overfitting such as Liang and Rakhlin, Bartlett et al. and others, which bound the excess risk with respect to the optimal hypothesis and also derive conditions on the data distribution for these results to hold. The results in this paper are not dependent on properties of the data, and the generalization guarantees are quite weak, hence I am not entirely convinced that model averaging is the reason behind the improved generalization, and it might be a red herring. Hence for the unconstrained case, I am not sure whether the model averaging interpretation offers the same understanding of generalization of overparameterized models as previous works such on benign overfitting. 2) Unless there is something I've missed in the paper, the model averaging properties seem to be properties of linear regression (and linear regression under simplex constraints) in general, and they are not specific to causal inference problems. Hence it might be useful to write the paper without the focus on causal inference, to make it appeal to a broader audience which might not be fluent with causal inference techniques. Instead, it would be nice to give the causal inference problems as examples/applications of the more generic result on the types of linear regression. 3) The paper does not discuss the interpretation of regression weights in the overparamterized case. Giving causal interpretations (under certain conditions) to regression weights is one of the major differences in using such models for causal inference, instead of for standard prediction tasks. Hence I'd expect some kind of discussion on such interpretations in this paper, as it focuses on causal estimation.
Questions
1) Are the results data-dependent in some way? Under what formal conditions can we give meaningful bounds in the residual error between the overparameterized solution and the optimal subset of donors (or w.r.t to some baseline donor selection method)? 2) Is there any point in trying to give a causal interpretation to regression coefficients in the method under some standard assumptions? While my intuition is that we should not do that, it is interesting to discuss this explicitly.
Rating
4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
4 excellent
Presentation
3 good
Contribution
2 fair
Limitations
The authors have properly discussed limitations, there does not seem to be a concern for negative societal impact.
Response to Rebuttal
Thank you very much for the rebuttal and your clarifications. To be precise about my current point of view on the paper, I think that the problem is interesting and I appreciate the treatment of overparameterization in the context of synthetic controls. I also think that the averaging property is something I did not know about and seems like a nice insight. However, without a strong generalization bound, it is unclear why model averaging as implied by the min-norm interpolator is beneficial for generalization. Concretely, we know there are cases where overfitting is not benign, hence model averaging can also be quite bad. I think that without a characterization of some intuitive data distributions where averaging helps us prove better generalization, the result is nice but lacks meaningful consequences. From reading proposition 5 I was not able to parse such a bound, since the average risk over all possible choices of donors seems like a pretty weak baseline, and I don't have any intuition about types of distributions where condition (P) holds (or even a toy example from which we can gain intuition). So while I am generally positive about this direction of work, the writing, and even like the current results, I think there are missing components that are important to the theory and its connection to practice. The additional experiments conducted with LASSO in response to reviewer xBLI are an interesting start in resolving such issues, but their connection to the theory is still weak and in my view (which may oppose the view of other people involved in the decision here), some more steps are required before publication.
Thank you for following up on our comments, and for providing further clarifications. We agree that additional theoretical results will be valuable to understand which lower-level conditions are sufficient to guarantee that convex model-averaging translates into improved performance from higher complexity. We hope that our manuscript already provides valuable contributions by (1) providing a tuning-free synthetic-control method that applies in the case of many control units, (2) documenting its properties on a real-world dataset, (3) deriving theoretical properties of the estimator, and (4) discussing how these results relate to interpolating regression. We believe that these contributions may motivate and feed into theoretical follow-up work along the lines you suggest. We believe that our results on synthetic control may also be of valuable practical relevance to applied researchers, when there is no strong prior on the importance of individual control units.
Thank you so much for the rebuttal and your detailed clarification. These answered my questions well!
Thank you very much for the rebuttal and your clarifications for the points raised by myself and other reviewers. I maintain my initial positive evaluation as it stands, albeit I somewhat agree with the concerns expressed by Reviewer TPer.
Thanks
I very much appreciate the rebuttal by the authors. I just wanted to re-iterate the importance of a more explicit study of the bias because Assumption (P) is a high-level sufficient condition. It would be good to mention this limitation in the camera-ready version. I changed my rating from 5 to 6 because the authors answered well most of my comments.
Decision
Accept (poster)