Non-Asymptotic Uncertainty Quantification in High-Dimensional Learning

Uncertainty quantification (UQ) is a crucial but challenging task in many high-dimensional regression or learning problems to increase the confidence of a given predictor. We develop a new data-driven approach for UQ in regression that applies both to classical regression approaches such as the LASSO as well as to neural networks. One of the most notable UQ techniques is the debiased LASSO, which modifies the LASSO to allow for the construction of asymptotic confidence intervals by decomposing the estimation error into a Gaussian and an asymptotically vanishing bias component. However, in real-world problems with finite-dimensional data, the bias term is often too significant to be neglected, resulting in overly narrow confidence intervals. Our work rigorously addresses this issue and derives a data-driven adjustment that corrects the confidence intervals for a large class of predictors by estimating the means and variances of the bias terms from training data, exploiting high-dimensional concentration phenomena. This gives rise to non-asymptotic confidence intervals, which can help avoid overestimating uncertainty in critical applications such as MRI diagnosis. Importantly, our analysis extends beyond sparse regression to data-driven predictors like neural networks, enhancing the reliability of model-based deep learning. Our findings bridge the gap between established theory and the practical applicability of such debiased methods.

Paper

References (93)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer HzkU3/10 · confidence 1/52024-06-25

Summary

This paper focuses on Uncertainty Quantification (UQ) in high-dimensional regression. The authors develop a new data-driven approach that applies both to classical optimization methods, such as the LASSO (which imposes an l1l1​ penalty on the weights), and to neural networks. They address the limitations of traditional UQ techniques like the debiased LASSO, which often produce overly narrow confidence intervals due to significant bias in finite-dimensional data. The authors derive non-asymptotic confidence intervals by estimating the means and variances of bias terms from training data, thus enhancing the reliability of confidence intervals for a large class of predictors.

Strengths

1. The paper seems to improve existing methods, though this is hard to tell (see weaknesses).

Weaknesses

1. The paper uses non-standard notation, making it difficult to read. In Theorem 1, $b$ seems to represent the target, which is typically denoted as $y$. Additionally, the relationship between the matrix A and the vectors $x$ is unclear. The l-1​ norm in equation (1) is applied to $x$, but $x$ is also referred to as IID data in Theorem 1. Generally, the l-1​ norm is used to penalize weights, commonly denoted by $w$, $\theta$, or $\beta$, rather than the input data for LASSO regression. 2. The paper does not clearly state the type of uncertainty being quantified, which could be clarified by addressing the first issue. 3. Some acronyms are not defined (e.g., MR, ITSA, LASSO). 4. Figure 1 is poorly presented. The images are very small with excessive white space in between, forcing readers to zoom in significantly. As a result, the caption becomes difficult to read.

Questions

See weaknesses.

Rating

3

Confidence

1

Soundness

2

Presentation

1

Contribution

2

Limitations

The authors do discuss the limitations of their method, but due to the lack of clarity in the text, it is difficult to assess these limitations effectively.

Reviewer WkBX8/10 · confidence 4/52024-07-12

Summary

This work develops an uncertainty quantification technique based on the debiased LASSO. The error is decomposed into noise and bias terms, which allows non-asymptotic confidence intervals to be derived. An empirical version of Chebyshev's inequality allows for their construction when the bias term is only assumed to have finite second moment, while sharper estimates are obtained in the setting where it is Gaussian. Numerical examples are given.

Strengths

This is a good paper and in my opinion should probably be accepted.

Weaknesses

The main weakness is that the proposed method is a competitor to conformal prediction, however there is no comparison of these methods or even mention of this. Some discussion on conformal prediction, and the relative merits of the new technique, are probably required for publication. The figures are too small, making them hard to interpret. This is compounded by the size of the text in the images. Their presentation should be rethought and fixed.

Questions

Can you please address the above issues?

Rating

8

Confidence

4

Soundness

4

Presentation

3

Contribution

4

Limitations

Limitations are discussed adequately. As mentioned, there is no mention of conformal prediction.

Reviewer BKX77/10 · confidence 3/52024-07-23

Summary

The paper presents a framework for constructing non-asymptotic confidence intervals around the debiased LASSO estimator. It derives a data-driven adjustment whereby the means and variances of the bias term of the debiased LASSO are estimated from the data and used to correct the confidence intervals. The framework is applied to the learned estimator from unrolled neural networks for real-world image reconstruction tasks, where the two moments are shown to be sufficient for modeling the bias term.

Strengths

- The non-asymptotic treatment is a promising and worthwhile extension to the debiased LASSO that's likely to benefit a variety of high-dimensional regression applications. - It's a convenient plug-in method around existing estimators of the debiased LASSO. - The experiments include representative settings where the remainder term is significant, and the relative norm $||R||/||W||$ is quantified for each experiment. - The coverage levels in the experiments are convincing overall, aside from a few remaining questions (see "Questions")

Weaknesses

See "Questions" for questions regarding the proofs and interpretation of experimental results. The text and figure formatting could be improved for clarity: - Please make the figures larger. The figures are missing axis labels and/or legends. Also, the tick and axis labels are too small. - Please label individual subfigures in addition to describing them in the figure caption (e.g., "(a) w/o data adjustment" for Figure 1). - For subfigures 1(d) and 1(e), and similar figures throughout the text, it would be helpful to overlay the confidence level $1-\alpha$ in a horizontal lines. - For Figure 3, please display (b) and (c) on the same y-axis scale. - L49: confusing phrasing, "when the dimensions of the problem grow" to describe the asymptotic setting

Questions

- For confidence intervals with significance level $\alpha$, the method often seems to have coverage beyond $1-\alpha$. Is the method prone to inefficiency or overcoverage of the CIs? It would be great to see some discussions in the experiments section as to where we could gain precision, for instance from the optimization of $\gamma$, and also refer to Section A in the main text. - What is meant by the "image support?" Could the authors please elaborate in general on what $S$ means and illustrate $S$ in the case of the MRI images for some selected $i$? - Why is $|W_j| \sim {\rm Rice}$ in L595 and not half-normal?

Rating

7

Confidence

3

Soundness

3

Presentation

2

Contribution

3

Limitations

The authors acknowledge that the accuracy of the method depends on the quality of the moment estimates and the ability to minimize the length over a larger parameter set, both of which depend on the data size. They also discuss opportunities to explore higher moments and other neural net architectures.

Reviewer WkBX2024-08-11

Response

I have read your response and it satisfies my concerns, particularly regarding the discussion on conformal prediction. I would expect to see this comparison in the main document, as it is key to the positioning of your contribution, as you have stated. I have raised my score accordingly.

Reviewer BKX72024-08-11

Thank you for addressing my comments. I will maintain my score. For the figures, the tick labels, the axis labels, and the red + markers should still be larger.

Program Chairsdecision2024-09-25

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC