Sub-optimality of the Naive Mean Field approximation for proportional high-dimensional Linear Regression

The Na\"ive Mean Field (NMF) approximation is widely employed in modern Machine Learning due to the huge computational gains it bestows on the statistician. Despite its popularity in practice, theoretical guarantees for high-dimensional problems are only available under strong structural assumptions (e.g., sparsity). Moreover, existing theory often does not explain empirical observations noted in the existing literature. In this paper, we take a step towards addressing these problems by deriving sharp asymptotic characterizations for the NMF approximation in high-dimensional linear regression. Our results apply to a wide class of natural priors and allow for model mismatch (i.e., the underlying statistical model can be different from the fitted model). We work under an \textit{iid} Gaussian design and the proportional asymptotic regime, where the number of features and the number of observations grow at a proportional rate. As a consequence of our asymptotic characterization, we establish two concrete corollaries: (a) we establish the inaccuracy of the NMF approximation for the log-normalizing constant in this regime, and (b) we provide theoretical results backing the empirical observation that the NMF approximation can be overconfident in terms of uncertainty quantification. Our results utilize recent advances in the theory of Gaussian comparison inequalities. To the best of our knowledge, this is the first application of these ideas to the analysis of Bayesian variational inference problems. Our theoretical results are corroborated by numerical experiments. Lastly, we believe our results can be generalized to non-Gaussian designs and provide empirical evidence to support it.

Paper

Similar papers

Peer review

Reviewer JMXP7/10 · confidence 3/52023-07-03

Summary

This paper derives precise asymptotic characterization for the naive mean field approximation in high-dimensional linear regression under the proportional limit. Two concrete corollaries are obtained including the inaccuracy of the NMF approximation for the log-normalizing constant and the overconfidence of the NMF approximation in terms of uncertainty quantification.

Strengths

1. The naive mean field approximation, the problem studied in the paper, is widely used in modern machine learning. The paper contributes to understanding the utility of NMF. 2. The paper is mathematically solid and all the technicalities are precisely presented. The proof techniques utilized are quite novel in the area of variational inference.

Weaknesses

I do not find any major weakness of the paper but there are some presentation issues that can be possibly improved. Since the paper is overall very technical and mathematical, I suggest the reviewers to setup a table or a subsection exclusively for all notations. Also, I suggest to considering to formalize conditions utilized throughout the paper such as Gaussian features, noises with special environments of assumptions instead of stating them within texts. The current presentation is a bit challenging for me to follow. Also, there are some typos that need to be fixed. For example, in line 78. I think the definition should be $G(u,d):=D_{KL}(\pi^{(h(u,d),d)}\Vert\pi^{(0,d)})=\cdots$; in line 129, the right most term should be $G(u,\frac{1}{\sigma^2})-\frac{u^2}{2\sigma^2}$.

Questions

1. In line 28, why the map $M_p(u)$ has to be defined over $[-1,1]^p$? It is a bit weird to me because $u=\mathbb{E}[\beta]$ but $\beta$ can be outside the interval. 2. The authors indeed adopt the parametrization based on exponential tilts $\pi^{(h(u,d),d)}$ on the basis of NMF. Does $\pi^{(h(u,d),d)}$ conceptually cover all candidates of continuous distributions? If not, is it the standard setup to facilitate analysis in variational inference?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

2 fair

Contribution

4 excellent

Limitations

See the weakness stated above.

Reviewer Qw445/10 · confidence 3/52023-07-05

Summary

The authors provide an asymptotic characterization of naive mean field (NMF) for linear regression under certain priors (e.g., those which guarantee strong convexity). Under these priors, this characterization can be used to demonstrate the sub-optimality of NMF (asymptotically). This is primarily a theoretical work, but some experimental observations of the fixed-point scheme are provided.

Strengths

The approach seems novel and fills gaps in the existing literature. The high-level ideas were easy to follow (though I didn't check all of the proofs in detail). The authors suggest lots of limitations which may lead to further investigation.

Weaknesses

It isn't immediately clear what practical utility this work would have. This isn't necessarily required in a theoretical work, but it would bolster the case for acceptance. Although the main ideas were novel (to me at least), the conditions required to draw the conclusions are difficult to assess (see questions below). The writing needs a lot of work. Some notation is undefined (though I could guess what it meant).

Questions

- Could you provide examples of possible practical applications of this work? The approach seems novel and interesting to me, but I can't ever recall using NMF and thinking that it would yield a highly accurate answer. - Could you comment on the conditions in Theorem 1 -- are these likely to be true in many instances? Let's suppose that neither condition is satisfied, is it possible to conclude anything from the fixed point iteration? Minor typos: - "i.e." and "e.g." must always be followed by commas. - "using product" -> "using a product" - "when feature" -> "when the feature" - "progresses were" -> "progress was" - Stylistically, it is better not to use a citation as a noun. You should use So and so et al. [citation] instead. ... too many more to log here. Serious proofreading is recommended, and please, please, please remove the phrase "not gonna be correct."

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

Yes, the authors clearly describe the limitations of the current work, which could suggest interesting avenues for future research.

Reviewer 6Ztj7/10 · confidence 3/52023-07-06

Summary

This paper studies the Naive Mean-Field Approximation (NMF) for the posterior distribution of a simple linear regression problem with Gaussian design. It provides an exact asymptotic result for the joint distribution between the NMF approximation $\hat u$ and the regression vector $\beta^\star$. This asymptotic is then used to provide results on the log-partition function associated to the NMF, the mean square error between the NMF estimator and the ground truth, as well as the associated confidence intervals around $\beta^\star$. Those asymptotics depend on a two-dimensional fixed-point equation, which can only be numerically solved; nevertheless, the authors compare this numerical solution with the full posterior distribution, to quantify how the NMF approximates specific quantities of interest of the posterior.

Strengths

This paper is an interesting contribution to the study of Naive Mean-Field estimators. It complements well the results of Sen and Mukherjee on the data-rich setting, by showing that the NMF approximation is not even correct to the first order in the proportional regime. The methods used are fairly standard in the field of exact asymptotics, which is acknowledged in the paper, but they are applied to obtain relevant insights on a new problem. As far as I checked, these methods are applied rigorously and the paper is correct.

Weaknesses

Some of the exposition feels however a bit too obscure: it feels like some very important points (e.g. the definition of the NMF objective and its explanation, or the main theorem) are drowned in very technical considerations (e.g. almost a full page of conditions to ensure the convexity of $F$, or the computation part of the CGMT). I think the paper could achieve a better balance between technical content and explanation, to reach communities that may be familiar with either exact asymptotics or Bayesian approximations, but not both. Some explanations that would maybe need to be highlighted include: - why you can replace (3) by the objective in Definition 2, and why you need the $\frac{d_i}2\beta_i^2$ correction term for, - how the CGMT also imply that the minimizers for both objectives are close (i.e. the part below Lemma 13) Minor remarks: - l.38: unfinished sentence ? - Eq. (3): $\sum_{i=1^p}$ - your log-normalizing constants are sometimes denoted with $Z$ (e.g. eq (2), (4)) and sometimes with $\mathcal{Z}$ (most of the time) - contrary to what is claimed l.203-207, the expressions for $\mathcal Z_p$ and related quantities are absent from the Supplementary Material - Eq. (9): I'm not sure redefining $F$ here is necessary - Fig. 5: there is no right panel

Questions

- How well do you expect the conditions of Theorem 1 to hold ? I read your Appendix C, so I understand why it would be complicated, but do you have at least a few examples of unicity ? - What's the use of Corollary 1 ? Isn't it strictly weaker than the conditions in l.120-124 ?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

N/A

Reviewer zux47/10 · confidence 4/52023-07-08

Summary

This paper studies the naive mean field approximation (NMF) in high dimensional linear regression in the limit where p/n goes to a constant. While the NMF is known to be consistent when n >>p, the authors show that in the high-dimensional case, this is not true. As a result, this paper gives nice theoretic evidence for empirical observations showing that the NMF is no longer accurate in high dimensional regimes. Moreover, the paper also discusses the consequences on the uncertainty quantification.

Strengths

The research question is important and the results should definitely be published at a top ML conference. The paper is well-written and easily understandable. The proof sketch is clear and the theoretical results are nicely illustrated with figures.

Weaknesses

While I recognize that this paper meets the standards for NeurIPS, my reluctance in awarding a higher grade stems from two significant limitations: 1) The proofs' methodological contribution is rather limited. Despite the elegant use of the CGMT, which is commonly seen within the high-dimensional statistics literature, the paper doesn't break new ground in terms of methodology. Although the paper commendably presents the proof in a digestible manner, and the choice of the CGMT is likely the most efficient way to prove the result, this doesn't warrant a significant methodological contribution deserving of a high grade (8+). 2) The insights gained from the results are fairly limited and expected. This is also generally a shortcoming of proofs relying on the CGMT. While such proofs usually yield the precise asymptotic limit, the resulting theorem statements are challenging to interpret.

Questions

I am currently missing an intuition in the paper for why the NMF approximation is loose in the high dimensional limit. Can the authors elaborate more on this point? To me, the main intuition arises from the fact that when $n>>p$, the eigenvalues of $1/n X'X$ concentrate around $1$, which is clearly not the case when $p/n \to c$. As a result, when $n>>p$, the measure $\mu$ is approximately symmetric around its mode (and thus a NMF approximation should be accurate). On the other hand, when $p/n \to c$, the measure $\mu$ is "elliptic" and strongly depends on the small eigenvectors of $X'X$. However, these eigenvectors cannot be captured by the NMF approximation.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Yes

Reviewer zux42023-08-10

Response to the rebuttal

I have read the rebuttal and the other reviewers comments. I recommend an acceptance of this paper. However, I strongly encourage the authors to focus more in their paper on the intuition why the NMF approximation is loose. I still think that the CGMT proof is fairly straight forward in this context and well exploited in related (frequentist) settings in the literature. Thus, I would rather put the proof sketch into the appendix and dedicate more space to a better reasoning why this behavior of the NMF is intuitive. Especially given that this is a conference paper, I believe that such a modification could strongly improve the impact of the paper. However, in the end this preference is subjective and may not be shared by other reviewers/the authors.

Reviewer 6Ztj2023-08-15

Thank you for your replies ! The outlined changes to the structure of the paper do alleviate my concerns about clarity, and hence I have changed my rating from 6 to 7.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC