Summary
The authors consider Shapley values computed from games $\nu$ modelled with a Gaussian process. Since the game is stochastic, the Shapley values are now stochastic, and the stochasticity reflects epistemic uncertainty in the explanations. This is unlike previous methods like BayesSHAP, where the uncertainty is estimation error. The approach of the authors, called GP-SHAP, can be used to explain predictions of Gaussian processes. It can also be used to construct a prior over Shapley values, which in turn can be used to perform regression on Shapley values. The authors demonstrate the utility of their approach is three experiments.
Strengths
I would like to thank the authors for their submission. I think this is interesting work.
## Strengths
* The paper is well written and mostly clear. The presentation is reasonably good.
* Pushing a GP through the value allocation computation to explain stochastic predictions seems like a very sensible thing to do.
* Doing the above to define a prior over value allocations and then using this prior for regression of Shapley values is, to the best of my knowledge, novel. I have not seen this before. I should, however, say that I'm not at all familiar with this part of the literature.
* The stochasticity in explanations by GP-SHAP represents epistemic uncertainty in the explanation, whereas the stochasticity in explanations by BayesSHAP is due to estimation error. I think this is an interesting finding.
* I have either skimmed through or looked in more detail at the proofs of the mathematical statements. Apart from some clarity issues (see below), I think everything generally checks out.
* The experiments nicely illustrate the benefits of GP-SHAP.
## General Remarks
* In Proposition 11, I was initially very confused by that the expectation of $v_x$ appears. After reading the proof, it became clear that this is an approximation that simplifies the computation, which should actually integrate over $f$. I think this can be better explained in the statement of Propostion 11.
* In Proposition 12, if I'm not mistaken, the transpose in the expression for $\kappa$ refers to taking inner products in the RKHS of $k$. This is _definitely_ not clear at all and should be clarified!
* In Figure 1, how does BayesSHAP produce its explanations? It is applied to the predictive mean of the GP prediction?
Weaknesses
## Weaknesses
* I think that providing a proof for Theorem 4 is unnecessarily complicating the exposition, as the Shapley's original proof for the deterministic case can be applied to immediately prove Theorem 4: Let $(\Omega, \mathcal{F}, \mathbb{P})$ be the underlying probability space. For any fixed sample $\omega \in \Omega$, $\nu$ is d-game (using the authors' language). Therefore, according to Shapley's original proof, for that $\omega \in \Omega$, it can be written in the form of (1). Since (1) holds for all $\omega \in \Omega$, it obviously holds as an equality in terms of random variables.
* Generally, I think that the exposition takes a long time to arrive at the key idea of the paper: model $\nu$ with a GP, and push it through the value allocation computation in (1). I think that the exposition would be improved by having Section 3 starting out with immediately explaning this key idea. A paper is not supposed to be a novel: please signpost important ideas and conclusions as much as possible.
* In lines 326-339, you highlight the difference between the mean of absolute SSVs and the absolute values of mean SSVs, and state that this gives a different ordering for which feature is most influential. However, you do not analyse whether this different ordering is better or worse, which means that it is not clear whether, in this case, GP-SHAP produced a better ordering or not.
* In lines 340-346, you produce a local explanation graphical model. However, you do not analyse whether this graphical model is reasonable, which means that the reader is not sure whether GP-SHAP did something sensible or not.
* Building a GP prior over Shapley values to do regression is an interesting idea. However, the presentation would be more convincing if you were to point out some existing applications that would actually want to do regression of Shapley values.
Questions
## Conclusion
I think this is a solid submission without any major shortcomings, which is why I am giving an accept.
If at all possible, I would like to see the following edits in a revision of the authors:
* If I'm right that Theorem 4 follows immediately from the deterministic case, I would really just do that and omit the current proof.
* Please start out Section 3 by explaning what you are working towards: pushing a GP through the value allocation computation
* Please clarify the statement of Proposition 11 (see above).
* Please clarify the transpose in Proposition 12 (see above).
* In Figure 2.(b), please analyse the difference in ordering between the mean of absolute SSVs and the absolute values of SSVs and conclude which of the two orderings is more sensible.
* In Figure 2.(d), please argue whether the graphical model is sensible or not.
EDIT
I have read the authors' rebuttal and the other reviews and remain at my current assessment.
Rating
6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.