Summary
The authors propose a new scheme to obtain prediction intervals that are valid conditioning on the model output. The approach leverages the notion of Perfectly Calibrated Point Predictors and an extension of Venn-Abers calibration to the regression setup.
Strengths
- I like the interpretation of $f(X)$ as a "scalar dimension reduction".
- Generalizing self-consistency to prediction sets may inspire future investigation in the broader CP community.
- The idea of leveraging Venn-Abers calibration to the regression setup is a valuable contribution.
Weaknesses
- The authors do not explain, in an intuitive way, why conditioning on $f(X)$ is a good approximation of conditioning on $X$. They could describe their 1D dimensional-reduction argument better, possibly with a concrete example. Imagine $(X, Y)\in {\mathbb R}$ and $Y \sim ({\bf 1}(X<0) + 10 {\bf 1}(X<0)) {\cal N}(0, 1)$. A perfect point-predictor is $f(X) = {\rm E}(Y|X) =0$ for all $X$. In this case, conditioning on $X$ or $f(X)$ does not look equivalent.
- The description of Venn-Abers calibration can be improved. I have not fully understood the meaning of these two sentences [1][2].
- The authors may explain more intuitively why they need a Perfectly Calibrated Point Prediction to compute a Self Calibrated Prediction Interval.
- The proposed model is only compared with Mondrian CP.
[1] *Venn-Abers calibration accounts for overfitting by widening the range of the multi-prediction in such scenarios, thus indicating greater uncertainty in the value of the perfectly calibrated point prediction.*
[2] *Each point prediction in the set enjoys the same large-sample calibration guarantees for isotonic calibration.*
Questions
- Intuitively, how does conditioning on $f(X)$ help? How does this avoid the *curse of dimensionality*? What is the difference compared with conditioning on $Y$?
- Is $1({Y \in C})$ the indicator of ${Y \in C}$? Why is $f$ called a *covariate shift* in (3)? How is this the same as saying that $f(X)$ is the model output?
- Has prediction-conditional validity been used before?
- How is $f^{X, y}$ defined? Should $f_{n+1} = ( f_n^{X_{n+1}, y}(X_{n+1}), y \in {\cal Y} ) $ be interpreted as a recursive definition? How does $f_{n+1}$ depend on $x$?
- Would it be possible to compute the marginal prediction band centered on the Venn-Abers calibrated predictor (the black lines in Figure 2, if I understand the plots correctly) instead of the original one?
Limitations
The authors should have said why they do not compare with other approximations of context-conditional validity, e.g. Conformal Quantile Regression, Ref 22 in the paper, or the Error Reweighting method [3].
[3]
Papadopoulos, Harris, Alex Gammerman, and Volodya Vovk. *Normalized nonconformity measures for regression conformal prediction.* Proceedings of the IASTED International Conference on Artificial Intelligence and Applications (AIA 2008). 2008.