Summary
The paper generalises the notion of kernel mean embeddings to higher-order cumulants. It proposes kernelled cumulants in the RKHS. While kernelled cumulants reside in tensor product space of the RKHS, the paper shows that Hilbert space metric between cumulants can be exactly computed using the kernel trick. Based on this construction, the paper proposes: (1) a two-sample test statistics, that generalises MMD test statistics by considering distance between cumulants, and (2) a generalisation of HSIC statistic for independence testing, again by considering distant between cumulants of joint distribution and product of marginals.
The advantages of the construction and proposed tests include: (i) the test are applicable for a broader class of "point-separating" kernels (unlike MMD/HSIC that can only useful for characteristic kernels); (ii) the new statistics can be computed in quadratic time (same as MMD); and (iii) empirically achieve higher power that classical MMD/HSIC statistics (both on synthetic and real data).
Strengths
- The idea of considering higher order moments/cumulants in RKHS is quite natural, and yet unexplored in the literature (apart from the recent work of Makigusa (2020) that consider only 2nd order moments). Hence, the contribution is novel and quite timely
- The main strength of the paper is that the computational cost of the proposed statistics is still quadratic in sample size (same as MMD), which implies the advantages of higher order moments does not come at significant additional cost
- The work is technically sound, and the construction of cumulant and use of kernel trick is reasonably involved.
Weaknesses
- While it is easy to imagine in general that higher-order cumulants can distinguish between more distributions, the advantage of kernelled cumulants is difficult to grasp. If a characteristic kernel is used wouldn't mean embedding (MMD) suffice?
- The line of argument used in the paper to demonstrate advantage of kernel cumulants is that (empirically) they have show higher power (can detect small differences better). Is there a theoretical justification for this? It would be sufficient if the authors provide a justification / reference that standard (non-kernel) cumulants are more sample efficient in some cases, where means already show separation
- For the synthetic experiments, the null rejection rate should also be plotted to show that cumulants do not have higher tendency to reject than MMD/HSIC. While this is true for the real data, unfortunately both 1st order and higher order terms reject at a rate higher than significance level
- Overall, it is not clear when tests based on higher-order cumulants are indeed needed in practice. I still feel the work is relevant, but some discussions on this would certainly increase the significance of the work for the broader community
- The paper, although well-written, is quite dense and at times bit difficult to follow, but this can be attributed to the content of the paper
Questions
- see weakness
- in addition, a precise statement on computational complexity (at least for d2, d3) would be useful
Minor remarks: H in Lemma 3 is not defined in main paper (but in appendix), and E et al citation seems incorrect (there is a full name for E)
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
the paper does not have immediate negative societal impact (although conclusions from hypothesis tests can always have). Hence, the work could benefit from:
- consistency results (similar to kernel two-sample tests)
- characterisation of whether higher-order cumulant based statistics typically tend to be larger than MMD (even under null). The comment is about sample estimates and not expected value (hence, more tied to concentration/consistency of test statistics)