Common statistical measures of uncertainty such as $p$-values and confidence intervals quantify the uncertainty due to sampling, that is, the uncertainty due to not observing the full population. However, sampling is not the only source of uncertainty. In practice, distributions change between locations and across time. This makes it difficult to gather knowledge that transfers across data sets. We propose a measure of instability that quantifies the distributional instability of a statistical parameter with respect to Kullback-Leibler divergence, that is, the sensitivity of the parameter under general distributional perturbations within a Kullback-Leibler divergence ball. In addition, we quantify the instability of parameters with respect to directional or variable-specific shifts. Measuring instability with respect to directional shifts can be used to detect the type of shifts a parameter is sensitive to. We discuss how such knowledge can inform data collection for improved estimation of statistical parameters under shifted distributions. We evaluate the performance of the proposed measure on real data and show that it can elucidate the distributional instability of a parameter with respect to certain shifts and can be used to improve estimation accuracy under shifted distributions.
Paper
Similar papers
Peer review
Summary
The paper presents a metric (s-value) that quantifies the uncertainty of statistical estimators in terms of their distributional instability. In addition, the techniques proposed can quantify the effect of directional shifts and the authors also discuss how the s-value can be used to improve estimation accuracy under shifted distributions.
Strengths
The development of methods that quantify the effect of distributions shifts is very relevant since such shifts are very common in practice. The overall approach in the paper is of interest and the results in the paper provide a promising initial step towards novel measures of uncertainty that account for distributional shifts.
Weaknesses
The interpretation of the s-value in (1) for scalar parameters is somehow clear since it measures the smallest divergence needed to change the sign of the parameter (assuming the parameter continuously depends on the distribution). However, the usefulness of its generalization to the multidimensional case in (11) needs further clarification. In the scalar case, the divergence needed to achieve a zero value reflects the divergence needed to have a change of sign. In the multidimensional case, it is not clear why measuring the divergence needed to achieve a zero vector valued \eta quantifies the instability. The experimental results are not correctly described. It is hard to grasp the main takeaways of such results and the relationship with the theoretical results provided. In general, the paper presentation is rather poor and multiple results appear in the appendices. It is clear that the paper can be significantly improved. The contribution of the proposed two-stage approach for transfer learning is hard to quantify. In general, the authors should compare their methods with existing techniques so that the paper's contributions can be better assessed.
Questions
The definition in (2) is unclear in page 2. The authors should mention there that the set \mathcal{P} in (2) denotes joint distributions of Z and E, since in that page those distributions correspond to r.v. Z only. In line 248 the authors mention "Our findings show that the average treatment effect is unstable with respect to changes in the marginal distribution of ‘age’, ‘education’, and ‘re75’. We also find that s-values conditional on ‘age’, ‘education’, ‘black’, ‘hispanic’, and ‘re75’ are non-zero, indicating that the average treatment effect can change its sign with a shift in the marginal distribution of these covariates (sX > 0.85)." Are those results shown in Figure 2? The authors should improve the quality of figures that currently is rather poor.
Rating
3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations.
Confidence
2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
2 fair
Presentation
1 poor
Contribution
2 fair
Limitations
The authors should improve the description of the limitations of the methods proposed as described in the "weaknesses" section.
Summary
This work defines a novel statistical measure of the “stability” of a parameter in a distribution with respect to changes in that distribution. From a high level, it is defined as the minimum KL distance to a perturbation that flips the sign of the parameter. A method is given to calculate the s-value for mean parameters and several experiments are performed to show its utility.
Strengths
The method extends to the directional case. I can see the application of this to average treatment effect to be useful in industry in online experimentation platforms. Strong arguments for the utility of s-value for doing robust transfer learning.
Weaknesses
One weakness is that the calculation of the s-value uses a theorem that is only applicable to mean value parameters. The difficulty is that the definition uses an optimization problem over a set of functions rather than real-valued variables.
Questions
Could you extend to more general cases by learning a flexible neural density approximator to P_0, and solving (3) by (constrained) variational inference?
Rating
3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
4 excellent
Presentation
3 good
Contribution
2 fair
Limitations
Even though it’s addressed in the appendix, it seems non-practical to calculate the s-value for parameters other than the mean or functions of the mean. The paper restricts discussion to distributions that are absolutely continuous with respect to the training distribution, although this is not much of a limitation as it covers most practical cases of interest.
Summary
This paper proposes a measure to quantify the instability of a statistical parameter under distribution shifts, which calculates the minimal KL divergence to flip the sign of the estimated parameter. The authors demonstrate its usage in helping to collect target samples in transfer learning. The idea is clear and the theoretical analysis is solid, while the experiments seem a little inadequate.
Strengths
1. The idea makes sense and the paper is easy to follow. 2. The proposed measure is novel, and the theoretical analysis is solid. 3. The proposed measure could help to collect data in transfer learning, and the two-stage method is simple but seems efficient.
Weaknesses
Generally, I love this paper, but there are some drawbacks that stop me from giving a higher score: 1. The experiments seem a litter inadequate. I view this paper as technical work, but there are almost no baselines to compare in experiments. There are some recent papers sharing similar ideas (efficiently collecting data in transfer learning), and I think the authors should compare or at least mention them. For example: * Data Shapley: Equitable Valuation of Data for Machine Learning. ICML 2019 * Shapley values for feature selection: The good, the bad, and the axioms * Algorithms to estimate Shapley value feature attributions * Towards Efficient Data Valuation Based on the Shapley Value Also, maybe the authors could compare with some feature selection methods in experiments. 2. There are some typesetting problems in Figure 4 on page 9. Some contents are covered by figures. 3. The authors measure the instability via the effects on estimation. I wonder whether it is a better way to directly measure the model performance since model performance has already taken into consideration of the estimated parameters. For example, this paper: "Minimax Optimal Estimation of Stability Under Distribution Shift Hongseok Namkoong, Yuanzhe Ma, Peter Glynn". I hope that the authors could discuss these two perspectives more.
Questions
Please refer to weaknesses.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
How could the proposed measure be used in neural networks or in large-scale problems?
Summary
For a given data distribution P0 and family of distributions around it, mathcalP, the authors propose a measure of stability in estimating a parameter theta(P) when there are distribution shifts within mathcalP. The idea is that you have data P0, but might really be interested in estimating theta for some P' != P0 that you might get when deploying your model. That is, from data P0, know how bad you will be at estimating theta(P'). Sensitivity to changes depend on the parameter of interest, and the idea here is to quantify how much of a shift it would take (within P, from P0) to change the sign of theta(P0). The resulting quantity s(theta, P0), which should maybe be caled s(theta, P0, mathcalP) is called the s-value, which takes values in [0,1] for 1 meaning "small shifts will change the sign of theta" to 0 meaning "the sign is stable". The authors not only study marginal shifts, but also those resulting when one knows that some given conditionals will be fixed. Finally, they describe a two stage procedure for determining if it would be useful to collect test set samples for certain subsets of features (those for which s value is most unstable).
Strengths
- distribution shift is discussed in ML for predictive tasks but less in statistical inference / parameter estimation settings. It's great that the authors are bringing together the two areas, and providing tools for reasoning about inferences in the real world where we do expect shifts - Moreover, we don't always care just about absolute shifts, but instead qualitative things, like parameter sign changes, which the authors focus on - The authors discuss the conditional case where shift is only limited to some features, which helps us study more specific instances of the problem as well as gain statistical efficiency when we do have some knowledge
Weaknesses
The weaknesses are not in the mathematical method nor experiments, but more so in the discussion and contextualization of this work. - I think the iterative two stage procedure for selecting features to collect test data for is a pretty cool use case, but there is some lacking discussion about when this would actually be feasible in the real world. Please further discuss when one should or should not be able to do this. - The authors somewhat quickly dismiss influence-function (IF) based estimation in the Related Work, but my strong feeling is that a more thorough discussion of the relationship to this field is necessary. IF based estimation is not just about robustness to outliers, but actually has a fairly deep connection to this work where influence functions show up in the functional derivative of a parameter to be estimated with respect to an underlying distribution. Moreover, just like the authors of this work consider the case of parameter sensitivity conditionally on certain factors of the joint distribution, the IF literature likewise considers projections that describe parameter sensitivity to changes in distributions when some factors are fixed. I admit that the IF literature is somewhat dense, but I believe it would be useful to state some relationships even at high level to the derivatives and Von-Mises ("distributional taylor") expansions in e.g. Kennedy's review here: https://arxiv.org/pdf/2203.06469.pdf ---- just some writing comments below --- - I assume figure 2 is about the data NSW given the legend mentioning demographics, but the figure did not explicitly mention this. The figure caption should mention the name of the dataset to avoid confusion. - Figure 3 does not mention which data it is about in the caption, and one has to work backwards to the NSW experiment to find out which data the figure is about by searching for the text "Figure 3". Please mention the data in the figure caption. - when you say "We employ Jin and Rothenhäusler [22]’s transfer procedure to estimate", it would be helpful to be able to read what the method is from a re-statement in your paper rather than having to open the citation, since it's a detail that is part of your experiments. - in lines 59-61, "We discuss how our ... same to re-estimate the parameter under shifted distribution", it would be correct to change "under shifted distribution" to either "under a shifted distribution" or "under shifted distributions". - The remark in 63-70 could be moved to the end of work since it sort of interrupts the flow
Questions
- The two-stage procedure involves testing the stability of various subsets of features and seeing if there are some for which it would help to collect samples from the test set. Since it would be usually not possible to do this kind of search in the real world, what would be the practical take-away for how to make use of this in non synthetic studies? Did I understand correctly that testing a given subset Xs requires collecting it from the test set? It's fine to study this ideal case in a research paper, but it seems like it necessitates more discussion. - NSW experiment: your initial estimator has a std dev of 492 which is large for an effect size of 820. How was the std dev estimated? Is this the expected ATE std dev using this method on this data? - There is a claim in the paper that "full transfer" doesn't provide much gain over "partial transfer". This probably needs some more explanation. Is the claim that using Pproj to estimate theta is as good as using Ptest? This must only be true under some assumptions. Even if these are mentioned implicitly when describing the data, it would be helpful for the authors to restate what might cause this phenomenon. It's fine to report this phenomenon if it does hold on your data, but please contextualize.
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
Yes.
Summary
This work introduces the s-value, a novel metric to measure the stability of statistical parameters. It is defined as the exponential of minus the largest KL-divergence for which the statistical parameter is 0. The smaller the s-value the more stable the parameter is. When the statistical parameter is the mean of a random variable, the s-value is shown to have a simple expression. The paper also introduces the less conservative directional s-value, which constrains the type of distribution shift possible.
Strengths
Many methods exist to find the worst case distribution shift within some radius as a way to obtain a robust estimator. This paper instead is finding the largest radius up to which a parameter is robust, and then use this value to derive a measure of stability. I find this approach novel and I can imagine this paper having some impact in its area. The idea behind the paper and the derivations are sound, I also appreciate the exhaustive appendix.
Weaknesses
I found the experiments hard to interpret, e.g. in figure 4 and 5, what is $\beta$? Maybe you could improve the legend for those figures to make them easier to interpret. In figure 4, why are the scales so different ([-2,2] vs [-200,100]). It might have been more valuable to dive into a single example and describe it more in depth. Considering the main focus of this paper, I would also suggest showing the s-values for different parameters e.g. in a table. Given the venue, I was expecting that the estimation of the s-values for parameters defined via risk minimization would be in the main paper. Overall, I appreciate the quality of this paper, but I would have liked to see more clearly how impactful the s-values can be in concrete cases. Achieving this might simply require improving the experiment section to address it to a broader audience.
Questions
How useful would s-values be for non-convex models? In the appendix, it is mentioned that a small s-value cannot be necessarily interpreted as a proof of stability in that case. What results would you expect to obtain if you were to run your parameter transfer experiment using a non-linear model?
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
2 fair
Contribution
3 good
Limitations
I think limitations could be better addressed by e.g. adding a paragraph in the appendix.
Summary
This paper proposes a novel metric, called the s-value, for evaluating the stability of statistical parameters with respect to distributional shifts. This metric is based on a variational problem involving the KL divergence between the target distribution and the shifted distribution, which can be solved via an equivalent one-dimensional convex problem. The authors also introduce the notion of directional s-value that quantifies the instability of directional shifts. Moreover, consistency and asymptotic normality results are proven for the plug-in estimators of these s-values. Finally, the authors illustrate the interest of the s-values on some real datasets.
Strengths
1. I found the proposal novel and interesting. 2. Quantifying the stability of statistical findings under distributional shifts is an important question and this work takes a first step towards answering this question. 3. The paper is well-organized and easy to follow.
Weaknesses
1. The treatment for the consistency of $\hat s_E(\mu, P_n)$ seems to be weak. In particular, (1) Estimating the conditional expectation $E[Z | E]$ is a challenging problem, especially when the dimension of $E$ is large as in Example 4. The uniform convergence assumption (Assumption 1 in Appendix C) seems to be too strong. (2) The rate of convergence of $\hat f_n(E)$ can be very slow when the dimension is large. It is of interest to know how the rate of convergence of $\hat s_E(\mu, P_n)$ depends on the one of $\hat f_n(E)$. 2. The practical usefulness of this metric is questionable. In the experiments the authors only computed a few s-values without giving empirical evidence supporting the reasonableness of these numbers. For example, is a statistical parameter with a large s-value really more sensitive to distribution shifts than one with a small s-value? It would be good to perform at least simulation studies to investigate this question.
Questions
1. Can you prove the consistency of $\hat s_E(\mu, P_n)$ under a weaker assumption on $\hat f_n(E)$? 2. How does the rate of convergence of $\hat s_E(\mu, P_n)$ depends on the one of $\hat f_n(E)$? 3. Can you provide empirical results supporting the practical usefulness of s-values? For example, is a statistical parameter with a large s-value really more sensitive to distribution shifts than one with a small s-value?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
3 good
Contribution
3 good
Limitations
Yes
Summary
This paper proposes method to quantify instability of a statistical parameter with respect to pertubrations around the KL divergence ball. This has implications to detect where statistical conclusions no longer hold when there is a distribution shift. The authors show this metric across both overall and directional shifted. The paper provides a two-step transfer learning strategy over two datasets to demonstrate the effectiveness of using stability as a measure of where to collect extra data for transfer learning.
Strengths
Tthe paper was well written, with its objectives and motivations clear. Theorems and mathematical notation were easy to follow. The examples in Section 3.2 illustrated well how this metric is concretely used for different distributions. It is evident that s-values can be useful in determining when to re-train models or re-estimate statistical queries, depending on the stability of a particular parameter. Detecting under which variables shifts occurs is an important open problem that the authors provide clear insight in, as well as a procedure to use s-values for improved transfer learning.
Weaknesses
However, while an interesting and intuitive idea, it does not seem that it provides significant improvements over a transfer learning approach with all covariates. In Figures 3 and 5, transfer learning outperforms the naive approach, but this is an expected result. What is the reason for a partial transfer if a full transfer works better than or just as well? I would also be interested to see an experiment where a greater amount of data can be collected to improve the transfer learning (aka a higher $\alpha$ value for the second experiment). It would be also helpful to provide recommendations for practitioners for what it means when a parameter is stable (is there a general threshold of $s$ at which transfer learning works well over a parameter works well?) Is the $s_X > 0.85$ threshold as shown in the paper recommended?
Questions
There are some uncertainties about the experimental methodology. Why were these particular datasets chosen? Something with an average treatment effect seems like it would be useful to be seen in medical datasets for a particular intervention? Would be interested to see if there would be significant improvements over standard transfer learning with all covariates if this procedure was done with datasets with more unstable covariate variables. It would be helpful to see a broader set of experiments, or a more thorough analysis of which covariates are helpful in this case or not.
Rating
5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
4 excellent
Contribution
2 fair
Limitations
Yes
I would like to thank the authors for their responses. However, I still think the paper contribution is not significant enough for this conference. For instance, it seems the methods proposed are only useful for the scalar case.
Misunderstanding
Thank you for engaging with our review. We believe that there has been a misunderstanding. Our procedure can be helpful in the multi-parameter case. For example, ANOVA can be used to test hypotheses about multiple parameters. ANOVA requires being specific about the null hypothesis (for example that the group means in ANOVA are all equal to zero). For such a hypothesis, one can also compute s-values as in equation (11) with eta=0. For more details, please see Appendix Section D. Thus, the proposed method applies to and can be useful in multi-dimensional situations as well, but scientific applications are usually formulated as inferences on scalar parameters in the presence of (potentially high-dimensional) nuisance parameters, which is the main focus of the paper. The focus on scalar parameters in the presence of nuisance parameters is a stylistic decision that reflects scientific practice. To the best of our knowledge, we are the first to study the stability of parameters under various types of distributional shifts.
Thank you for your detailed explanation. I updated my score to believe the following will be reflected in your revised manuscript (or camera-ready version). * A discussion on relevant works measuring stability, e.g. “Minimax Optimal Estimation of Stability Under Distribution Shift”. * A discussion on the choice of KL divergence, especially its potential shortcomings.
Author's response to your concerns
Dear reviewers, The authors have responded to your questions and reviews. Since the discussion is coming to an end, we are wondering if you could kindly take a look and see whether the responses address your concerns and whether you'd like to update/maintain your initial rating. Best, AC
Decision
Accept (poster)