Summary
This paper studies the learning theory aspect of the domain adaptation problem, where the key is to bound the estimation errors between expectations over shifting distributions. Specifically, this work improves the recently developed $f$-divergence-based generalization analysis, where the main results ensure a tighter generalization upper bound and the consistency between theory and method. For finite sample setting, a sharp bound is provided to accelerate the asymptotic rate. Numerical simulation is conducted to demonstrate the superiority of the theory-guided method over the existing discrepancy-based framework.
Strengths
+ The motivation is clear, i.e., improving the $f$-divergence-based bound and bridging the gap between method and theory, and the presentation is easy to follow.
+ The technical part is generally sound and the justifications are sufficient.
+ The experiment results are superior compared with recently developed generalization bounds.
Weaknesses
+ Some notations are inconsistent in theoretical analysis.
+ The proposed algorithm needs further justifications.
+ The experiment comparison could be improved.
Questions
There seem no major faults in this submission, and I only have the following minor concerns.
Q1. Theory and methodology. The major result for the target error bound is provided in Eq. (4) in Thm. 4.1 and the specific bound w.r.t. KL-divergence is presented in line 162, where the induced learning objective consists of source risk and the square root of cross-domain KL-divergence. However, it seems that the optimization objective Eq. (5) considers the divergence without the square root directly. I understand the optimal solutions are the same for these two objectives (if the optimal solutions ensure 0 cross-domain discrepancies). But considering Eq. (4) is closely related to the major merit of this work, i.e., the tight bound, the consistency between Eq. (4) and Eq. (5) seems to be important. Some justifications are highly expected.
Q2. Method application. As far as I understand this work, the derived $f$-DD measure can be applied to existing works whose primary goal is discrepancy minimization. Thus, it could serve as a plug-and-play module for existing SOTA DA methods. Thus, some detailed discussions on the capability of $f$-DD w.r.t. existing methods are highly expected.
Q3. Following Q2, apart from the experiments in the current version, some comparisons between SOTA DA methods and their combination with $f$-DD objective are highly expected.
Q4. The clarity w.r.t. definitions could be improved, e.g., $K_{h',\mu}(t)$ depends on the hypothesis $h$ while the justification (i.e., line 178) is provided after the definition (i.e., line 176). A thorough check for these issues could improve the readability.
Q5. Some notations seem to be inconsistent. For example, the notation $I_{\nu}^{\phi}(h,h')$ in line 132 is inconsistent with $I$ in line 129; the notations $\mathbb{E}_{\nu}$ in line 132 seems to be incorrect (probably should be expectation over $\mu$?).
Limitations
The limitations are discussed in the checklist, and there seems no potential negative societal impact.