Summary
This paper studies the theoretical aspect of gradual domain adaptation (GDA), where the knowledge of labeled source domains is supposed to be transferred to a sequence of target domains. The main results show that the gradual adaptation process can be well characterized by the distributionally robust optimization (DRO) framework, where the domain gaps between the source domain and multiple target domains are gradually captured by the robustness of the model within a pre-set region, i.e., a ball w.r.t. Wasserstein metric over probability space. Finally, the distribution shift is guaranteed to be mitigated with the DRO algorithm.
Strengths
+ The motivation of employing DRO as a theoretical framework to address gradual domain adaptation is clear and reasonable.
+ The extensive theoretical results seem to be solid.
Weaknesses
- The clarity should be improved, where the advantages or improvements w.r.t. related theory for GDA is not discussed, which leads to unclear contributions.
- The presentation should be improved, e.g., the DRO algorithm is provided without justification while the main results closely depend on this algorithm.
- There are many typos and the readability is fair.
Questions
Q1. As discussed in the previous work section, there are already several theoretical works for GDA, e.g., [WLZ22] and [HWLZ23]. However, the differences between the derived results and these works are not discussed in either *previous work section* or *main result section*. It should be clarified that what new insights are provided in the derived results.
Q2. The basic framework, i.e., DRO, is presented in Algorithm 1 directly, while there are no insights provided for it. Though there are plenty of results derived based on DRO, it is hard to understand the working mechanism of the DRO algorithm.
Q3. The bond in Theorem 2.3 shows that the target error can be dominated by the source risk with the factor $g_\lambda (\cdot)^{\circ T}$. Though Corollary 2.4 provides an analytic bound for $g_\lambda$, it would be more interesting to show the monotonicity of $g$ w.r.t. the composition operator. Furthermore, can the factor $g_\lambda (\cdot)^{\circ T}$ can be monotonically reduced by the increase of $T$? i.e., the factor of error can be reduced with the gradual adaptation process.
Q4. In literature [WLZ22], the main result of Theorem 1 shows that the target error is bounded by the source risk and accumulated error w.r.t. domain number $T$, which is induced by gradual adaptation process. Note that the main result in submission show the accumulated error is a factor of source risk, which implies this result could be loose when the factor is large. Some comparison between the tightness of these bounds are highly appreciated.
Minor: 1) Line 126, reference of Theorem; 2) Line 140, notations $\theta$ and $\Theta$ are used without definitions; 3) Line 186, previous studies [].
Limitations
The theoretical analysis for gradual domain adaptation seems to have no potential negative societal impact.