Distribution-Free Model-Agnostic Regression Calibration via Nonparametric Methods

In this paper, we consider the uncertainty quantification problem for regression models. Specifically, we consider an individual calibration objective for characterizing the quantiles of the prediction model. While such an objective is well-motivated from downstream tasks such as newsvendor cost, the existing methods have been largely heuristic and lack of statistical guarantee in terms of individual calibration. We show via simple examples that the existing methods focusing on population-level calibration guarantees such as average calibration or sharpness can lead to harmful and unexpected results. We propose simple nonparametric calibration methods that are agnostic of the underlying prediction model and enjoy both computational efficiency and statistical consistency. Our approach enables a better understanding of the possibility of individual calibration, and we establish matching upper and lower bounds for the calibration error of our proposed methods. Technically, our analysis combines the nonparametric analysis with a covering number argument for parametric analysis, which advances the existing theoretical analyses in the literature of nonparametric density estimation and quantile bandit problems. Importantly, the nonparametric perspective sheds new theoretical insights into regression calibration in terms of the curse of dimensionality and reconciles the existing results on the impossibility of individual calibration. To our knowledge, we make the first effort to reach both individual calibration and finite-sample guarantee with minimal assumptions in terms of conformal prediction. Numerical experiments show the advantage of such a simple approach under various metrics, and also under covariates shift. We hope our work provides a simple benchmark and a starting point of theoretical ground for future research on regression calibration.

Paper

References (69)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer bxwF4/10 · confidence 5/52023-07-05

Summary

This paper addressed the uncertainty quantification problem in regression models, specifically focusing on individual calibration to characterize prediction model quantiles. To overcome these limitations, they proposed simple nonparametric calibration methods that are both computationally efficient and statistically consistent. Their approach provides insights into individual calibration possibilities and establishes upper and lower bounds for calibration error. They attempted to advance existing theoretical analyses by combining nonparametric and parametric techniques, offering new perspectives on regression calibration regarding the curse of dimensionality and reconciling previous findings on individual calibration impossibility.

Strengths

The paper is overall well written and the guarantees given are technically sound. The authors have demonstrated the effectiveness of their methods through several experiments. The ideas developed are novel to the best of my knowledge but their practicality is questionable. There are many advancements in the field of conformal inference that give similar guarantees with minimal assumptions. I think the paper would benefit from comparing their developed method in varied settings with the existing approaches in conformal inference to prove its efficacy.

Weaknesses

In this paper, the authors tackle the uncertainty quantification problem for regression models, focusing on individual calibration. While the proposed nonparametric calibration methods and the accompanying analysis present some interesting ideas, I have serious concerns about the authors' familiarity with previous works in the field, as well as the lack of necessary citations to support their claims. One significant issue is the absence of references to previous research on conditional conformal prediction, which is a relevant and well-established framework in uncertainty quantification. The authors should have acknowledged and discussed how their proposed methods relate to or differ from existing conditional conformal prediction approaches. Failure to do so undermines the paper's novelty and raises doubts about the authors' understanding of the current state-of-the-art in the field. There have been various works in this area namely: 1. "Conformalized Quantile Regression" by "Yaniv Romano, Evan Patterson, Emmanuel J. Candès", 2. "Improving conditional coverage via orthogonal quantile regression" by "S Feldman, S Bates, Y Romano" 3. "Class-Conditional Conformal Prediction With Many Classes " by "T Ding, AN Angelopoulos, S Bates, MI Jordan, RJ Tibshirani" 4. "Conformal prediction with conditional guarantees" by "I. Gibbs, J. J. Cherian, E. J. Candès " 5. "Knowing what you know: valid and validated confidence sets in multiclass and multilabel prediction" by "M Cauchois, S Gupta, JC Duchi" The proposed method also suffers from the curse of dimensionality and the authors have only included experiments with several UCI datasets which are relatively low-dimensional. Whereas on the other hand, advancement in the field of conformal inference has produced methods that seamlessly adapt to very high dimensional datasets like ImageNet, MNIST, etc. Further, the theoretical guarantees in the paper rely on very strong assumptions, i.e., Assumption 1 of the paper, unlike previous works in this field.

Questions

I do not have further questions, all the concerns are summarised in the "weakness" and "limitation" sections.

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

1 poor

Limitations

The proposed method suffers from the curse of dimensionality as I mentioned before and the method relies on very strong assumptions. The authors have not implemented their proposed method on high dimensional datasets.

Reviewer 2q3v6/10 · confidence 3/52023-07-07

Summary

This paper studies uncertainty quantification for the regression problem. In particular, it considers the estimation of conditional quantiles (of the residuals) via the kernel method. The convergence rate of the proposed estimator is established, along with a matching lower bound. The proposed method is evaluated on multiple datasets and compared with other candidate methods.

Strengths

The paper considers an interesting problem; the examples for showing the unexpected results of existing methods are motivating; the solution provided has solid theoretical properties and show satisfactory empirical performance in numerical experiments.

Weaknesses

I was wondering about the position of this paper in the line of works of conditional quantile regression (e.g., Takeuchi et al. (2006); Steinwart et al. (2011)). A discussion in this direction will be appreciated. References: Takeuchi, Ichiro, et al. "Nonparametric quantile estimation." (2006). Steinwart, Ingo, and Andreas Christmann. "Estimating conditional quantiles with the help of the pinball loss." (2011): 211-225.

Questions

1. As mentioned above, I wonder how this work compares with the line of works of (conditional) quantile regression. 2. It might be helpful to also show in the simulations the results if one directly estimates the conditional quantiles of $Y$. 3. I wonder if the proposed method can be used within the framework of conformal inference and achieve distribution-free marginal calibration as well.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors have adequately addressed the limitations.

Reviewer oVNP6/10 · confidence 3/52023-07-08

Summary

The paper considers the uncertainty quantification problem for regression models. First, they proposed an algorithm for simple nonparametric quantile estimator. Then, they further proposed the nonparametric regression calibration algorithm. They also provide theoretical analysis of the proposed algorithms and implications.

Strengths

1) The paper is well-organized and well-written. 2) The paper provides several theoretical results and implications of the theory. 3) The paper includes extensive experiments.

Weaknesses

1) For Algorithm 1, do other kernels work for the proposed algorithm? Are there any requirements for the kernels? 2) For Algorithm 2, could the split proportion be different than half and half? The proportion would influence the results. Are there any experimental results to check the effect of the proportion?

Questions

1) For Algorithm 1, do other kernels work for the proposed algorithm? Are there any requirements for the kernels? 2) For Algorithm 2, could the split proportion be different than half and half? The proportion would influence the results. Are there any experimental results to check the effect of the proportion?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

NA

Reviewer aPS18/10 · confidence 2/52023-08-01

Summary

This paper proposes a new method for quantile regression and for calibrating prediction intervals in regression. The paper first proposes a simple quantile regression method and shows that this method estimates the true quantile curve at the minimax-optimal rate. For calibrating prediction intervals, the basic idea is to 1) decompose the conditional distribution of the response variable into a conditional mean + noise term, 2) estimate the conditional mean, and 3) apply the above quantile regression method to the distribution of residuals (i.e., respose minus estimated conditional mean). The approach is agnostic to the method used to fit the conditional mean itself, and the experiments demonstrate that the proposed method (as well as a supplementary method that also applies dimension reduction) performs well using feed-forward and recurrent neural nets, as well as random forests.

Strengths

The paper is quite easy to follow. The proposed method seems both simple and effective, and makes very weak assumptions. Figure 1 makes the advantage of individual calibration very clear, and the theoretical guarantees are hence both useful and impressive in light of existing results on the impossibility of individual calibration. The experiments are also fairly thorough and well-presented.

Weaknesses

1) Table 1: The "Ours Best?" column seems misleading, because it is taking the best of two different methods (NRC and NRC-DR). In several cases, only one of the proposed methods performs best, while the other method performs worse than competitors. However, in practice, one must typically pick one method before knowing which of the two will perform better. So, this column is not really informative of how the proposed methods would perform in practice. I suggest either removing this column. Perhaps a vertical rule could be added to distinguish the current paper's methods from previous methods. 2) The introduction is quite long, and it is not clear to me whether all of this information should be presented so early in the paper. For example, the discussion on Page 2 about "individual calibration" is unclear since "individual calibration" isn't defined until Page 3. Much of the content could also be refactored into a specific "Related Work" section.

Questions

1) One of the paper's key ideas is to decompose the prediction interval problem into a mean regression problem and a noise quantile prediction problem. Are there any cases where this decomposition would fail, or would be expected to perform worse than a method that learns the regression quantiles directly? 2) Relatedly, how does the performance of the mean estimator affect the performance of the prediction interval (e.g., if the mean estimator under- or overfits)?

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

The paper is fairly clear about its limitations, such as the gap between the theoretical performance (which suffers from the curse of dimensionality) and the strong empirical results.

Reviewer bxwF2023-08-14

We thank the authors for the clarifications. "Also, as noted in our paper, we found researchers from different communities, conformal prediction, calibration, and quantile regression, working on similar or even identical problem setup but don’t acknowledge much each other’s work."-- I am sorry but I am not convinced by these line of reasonings. I believe we are posing a problem relevant to the scientific community and proposing a solution that works and somewhat better than existing methods (thus, adding to the novelty). The field of conformal inference has flourished in the recent times and has shown exceptional promise in solving the posed problem both in regression and classification settings. As for split conformal, we have methods like Jackknife + (Barber et al. 2022) which seamlessly avoids data splitting. Given the generalizability of conformal inference type methods across different dimensionality of datasets, types of problems (regression and classification), and computational efficiency, the paper stands incomplete without proper comparison.

Authorsrebuttal2023-08-15

We thank the reviewer for the follow-up. In writing the paper and the previous responses/discussions, we try to follow two principles: (i) open-mindedness; (ii) technically rigorous arguments. We understand that this can be a busy week for all the reviewers, but we appreciate it very much if the reviewer can spend some time reading our paper and the previous responses. Our apologies for writing a lengthy response as before. For the sake of saving time, we provide a TL;DR version for the literature mentioned by the reviewer below: - Paper 3 and Paper 4 are posted on ArXiv after the NeurIPS 2023 submission deadline. We are not sure if these two papers are also under review at NeurIPS 2023. - Even so, we aim for open-mindedness and will include these two papers in future versions of our paper. In our previous response, we discussed that Paper 3 and Paper 4 are not aiming for an individual coverage guarantee, which differs from our positioning of individual calibration. - Paper 1 and Paper 5, together with the new paper Jackknife + (Barber et al. 2022), mentioned by the reviewer, all aim for a marginal calibration guarantee. For the newly mentioned paper (Barber et al. 2022) apart from the marginal calibration guarantee v.s. our individual calibration guarantee, the proposed method can be computationally costly compared to our method because it is a Jackknife-based approach. - Paper 2 aims for individual calibration but requires much stronger conditions for their theoretical results (See the last few paragraphs in our Author rebuttal). We see all above as “facts” that are supported by technically rigorous arguments. We are happy to follow up with additional technical discussions if there are any confusions/comments with these facts. We’d like to separate “facts” from “opinions”, such as the following one which is disagreed by the reviewer: "Also, as noted in our paper, we found researchers from different communities, conformal prediction, calibration, and quantile regression, working on similar or even identical problem setup but don’t acknowledge much each other’s work." This is our opinion, which of course, can be agreed or disagreed with by other researchers. We have such an impression/opinion because of the fact that for the 50+ papers in our reference list, the calibration papers mostly don’t mention the work on conformal prediction, while the conformal prediction papers don’t cite the calibration papers. However, even before this discussion, we try to maintain an open-mindedness and draw connections between our results and the conformal prediction literature in our paper. We thank all the reviewers and ACs for taking the time to read our responses.

Reviewer bxwF2023-08-17

I appreciate the authors for their follow up. However I maintain my rating for the following reasons: 1) "Paper 3 and Paper 4 are posted on ArXiv after the NeurIPS 2023 submission deadline." Yes and in that case you do not need to cite these works. My intention to suggest a list of few papers ignoring many others was to draw attention of the authors that active research on similar problem is undergoing and some of them should be added as a baseline. 2) Jackknife + (Barber et al. 2022): This paper has faster implementation (see their K-fold cross validation version) which drastically improves computational efficiency. 3) Indeed finite sample conditional coverage is impossible without further assumptions but several attempts have been made in the field of conformal inference. As mentioned in Author's rebuttal above, paper 2 comes close to their paper and their is no discussion or benchmarking in experiments with the submitted paper. 4) An essential focal point within this domain pertains to the size of sets. Extensive set sizes tend to lack informative value, constituting a significant metric frequently employed for appraising the effectiveness of prevalent predictive inference techniques highlighted across numerous conformal inference publications. Regrettably, this metric is absent in the current paper. While I acknowledge potential constraints such as the absence of cross-community citations, addressing these omissions and conducting thorough benchmarking would undoubtedly elevate the scholarly contribution. 5) Paper 2 works under a different set of assumptions but is easily adaptable to a wide range of real datasets and particularly high dimensional set up as established in their experimental set up. While the proposed method in the submitted paper again relies on an underlying low dimensional structure for a high dimensional feature set. I am finding it hard to believe the scope, applicability of the proposed method, and efficacy in terms of set sizes in wide range of real datasets. I again appreciate the authors for their work and I think the paper would strongly benefit with considerations of relevant work from the field of conformal inference and validate the novelty of contributions particularly in real data because these methods are widely used in many field of applications and have shown promising results.

Authorsrebuttal2023-08-21

We thank the reviewer for the helpful comments. We agree with the reviewer and appreciate the conformal prediction community for their efforts and results, and we find the papers mentioned very inspiring. In the past few days, we implemented numerical experiments to compare our algorithm with the algorithm in Paper 2 on the 6 datasets that are of the highest dimensions among the datasets in Paper 2. Such conformal prediction benchmarks and detailed experiments will be included in our further versions. 1. We apologize for missing some of the conformal prediction literature, but indeed we tried our best for the literature review. We are willing to add more discussions as mentioned in our previous responses in the future. Among the conformal prediction literature, we aim to address the minimum assumption to keep both the individual calibration (or coverage) and the finite sample guarantee, reconciling the impossibility results for quantile calibration (see our Appendix B) as well as the impossible triangle in the field of conformal prediction. Theoretically, we don’t find any paper in the literature on conformal prediction that gives a better result with comparable or weaker assumptions (including the papers mentioned by the reviewer). We refer to the previous Author Rebuttal for details. 2. As for the (practical) performance metrics, although the quantile calibration task is very closely related to the conformal prediction task, the goals are still different. Conformal prediction aims to provide as sharp as possible prediction sets that cover the true outcomes with desired rates. While a precise quantile prediction guarantees the coverage rate, such a covering band may not be the sharpest. As is discussed in Appendix A, the additional sharpness regularization term could harm the goal of precise quantile prediction. For example, people may select 0.05 and 0.95 quantiles to construct a 0.9 coverage band, but the sharpest covering band probably would deviate from such quantiles. But a high-quality quantile prediction alone is of independent interest for many downstream tasks, such as the newsvendor problem. This points to a difference between the objectives of calibration and conformal prediction. As a result, the popular performance metrics applied in conformal prediction such as the average interval length and the coverage rate are not direct measurements for the quantile calibration problem. Some other measurements, for example, the independence between the coverage indicator and the band length in Paper 2, are only a necessary condition of the quantile calibration. We henceforth consider only the explicit measurements rather than such implicit ones in our paper. 3. Numerically, to further illustrate the issue of high-dimensionality, we selected the 6 highest-dimensional datasets that appear in Paper 2 and applied our algorithm and Paper 2’s Pearson-correlation-regularized algorithm. First, we note that the first step regression of our algorithm doesn’t have to be the mean regression but can be any regression algorithm. We apply a quantile regression algorithm as the first step, and then recalibrate the initial result using our nonparametric estimator. As for the metric, we select (a 90% confidence interval variant of) the Adversarial Group Calibration Error (AGCE, defined in Appendix D.1) to show calibration error in the worst calibrated part of the data. We list the results as follows (Paper 2’s OQR the former, and the latter ours): meps_19: 0.071, 0.03 meps_20: 0.059, 0.031 meps_21: 0.043, 0.036 facebook_1: 0.041, 0.022 facebook_2: 0.011, 0.018 blog_data: 0.03, 0.024. The numerical performance shows the advantage of our algorithm. Of course, AGCE is a conditional/individual performance measure; for marginal calibration objectives or other datasets, conformal prediction methods may have an advantage. Just like the general ML problem, we don’t expect a single ML model that performs universally well, but we do believe our algorithm provides a simple and efficient complement to the existing methods. 4. As for the high-dimensional cases, we here make some further explanations. On one hand, the belief that high-dim data can be represented by lower-dim features is common in modern machine learning. Both computer scientists and statisticians are making efforts to extract low-dim features from the original high-dim datasets. On the other hand, the tremendous impact of reducing the Lipschitz coefficient $L$ partly explains why our algorithm works even in the presence of high dimensionality (and also why there is a splitting procedure in the split conformal prediction). As mentioned by the reviewer, there are conformal prediction algorithms that don’t follow the split protocol, but to our knowledge, we provide a first explanation/justification for those split conformal prediction algorithms. We thank the reviewer again for their helpful comments, which we believe consolidate the overall positioning of our paper.

Reviewer aPS12023-08-16

Thanks to the authors for their response. Regarding the decomposition: I understand why this approach reduces the Lipschitz constant of the estimand and thereby reduces the difficulty of estimation. My question was whether there are any counterexamples where this would be expected to perform poorly (or less well than a one-step procedure). If so, the paper would be made stronger and clearer by discussing such a counterexample. Regarding the writing of the introduction, I understand the reason for the length (and I appreciated what felt to me like a thorough literature review). My suggestion was that it could be better organized for the reader (e.g., by using some more subsection/paragraph headers, or by moving some of the content that isn't needed to understand the proposed method into a "Related Work" section). Although it's important to point out gaps in the existing literature (e.g., in a "Related Work" section), I feel that the basic motivation for proposing an approach shouldn't depend so heavily on the existing literature (which is always changing). I don't necessarily expect the authors to reply to the above points in the discussion period -- just some things to think about as they continue to revise the paper. Overall, I still feel this paper gives a simple, effective, and novel approach to an important problem, backed by solid experiments and theory, and so I intend to keep my score of 8. That said, I am not up-to-date on the conformal inference literature (the main reason for my low confidence of 2), and I defer to Reviewer bxwF on whether critical papers are missing in this regard (beyond what the authors could easily add to the camera-ready version).

Authorsrebuttal2023-08-21

We thank the reviewer for raising the points; we will continue thinking about them and include more content in these aspects in the future version of our paper.

Reviewer 2q3v2023-08-20

Response to the authors

I thank the authors for the response and the comparison, and I look forward to seeing these contents added to the paper.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC