Summary
This manuscript presents convergence rates for kernel methods under covariate shift. Results fit quite a general framework, including common classification and regression losses. Two approaches are analyzed: (i) a usual M-estimator and (ii) an importance-sampling-like M-estimator. It is shown theoretically and empirically that the latter outperform the former.
Strengths
The analysis presented in this paper provides interesting theoretical results regarding learning under covariate shift, which is a contemporary topic. The manuscript is well organized; it explains clearly the problem, state the results while discussing the hypotheses and, at the end, illustrates the theoretical findings by a numerical experiment.
I would like to stress that discussions regarding hypotheses are opportune and corollaries provide intelligible results.
The take-home message, stating that the importance-sampling-like estimator is better that the naive one, is interesting and confirms practitioners’ intuition.
Weaknesses
Major remarks:
1) My main concern is about the novelty of the proofs: hypotheses (i) and (ii) look like straightforward tools to link expectations under the source distribution to the target distribution by linearity or Cauchy-Schwarz inequality. I had a very quick glance to the supplementary material and it confirmed this guess (although I admit that I may be wrong). I think that its important, in order to assess the contribution of the paper, that the authors explain the original derivations appearing in the proofs, with respect to techniques used for obtaining similar results without covariate shift (unfortunately, I have no reference in mind).
2) Another (minor) point is that Figure 1 does not seem to verify neither hypothesis (i) nor (ii) since $\phi(x)$ seems to explode when $x \to \infty$. If it is the case, it would be better to find another example (or at least to discuss this point). If it is not the case, it would be informative to explain it.
Some suggestions of improvement:
1) $f^*$ is defined in Section 2.1, before the problem setting in Section 2.2. However, in practice, it corresponds to the optimal function under the target distribution, which is not clearly stated. I suggest to make it clear.
2) Although an informed reader understand definitions Line 113, it is not totally clear that expectations are conditioned by observed data. I suggest to had this information.
3) $D$ could be added after “Finite rank” in Table 1.
4) Line 264, it is not totally clear that “For the moment bounded case” correspond to Figure 3. I suggest to had it.
Typographical remarks:
1) Extra “the” Line 5.
2) “that” instead of “that is” Lines 126, 131 and 132.
3) In Theorems 1-3, $\delta_n$ should satisfy an inequality that involves $\delta$ instead of $\delta_n$.
4) There are $\phi(\textrm x)$ instead of $\phi(\textrm x)^2$ Lines 212 and 223.
5) Full points are missing in captions.
Questions
1) What is $f_j$ Line 154?
2) What does $\psi_j \le Cj^{-2r}$ (Lines 201 and 235) mean, given that $\psi_j$ is a function? Is a norm missing?
3) Does the trends (evolution with respect to $n$) observed in Figure 2 (b)-(e) and Figure 3 (b)-(e) is of the order $\left( \frac{\log n}{n} \right)^q$ or $\left( \frac{\log^2 n}{n} \right)^q$ as exhibited in Table 1?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
Limitations are not addressed.