Summary
This paper attempts to build the implicit regularization of gradient flow (GF) for the deep diagonal linear networks. In particular, the authors show that the GF dynamics induces a certain kind of mirror flow dynamics without an explicit form of the corresponding entropy function. In addition, they also reveal the convergence property of the GF dynamics and establish how the convergence rate is affected by the initialization.
Strengths
In general, this paper is well organized, e.g., the authors clearly demonstrate their motivation and contribution. They also clearly develop their notations, definitions, and theorems to support their claims. These efforts make the understanding of this paper fairly straightforward. In addition, the characterization of certain properties of the learning dynamics of deep diagonal linear networks might also be interesting, e.g., the second point of Proposition 1.
Weaknesses
Unfortunately, both the technical and theoretical contributions of this paper are rather limited, which I will discuss as follows.
1. The ultimate goal of this paper is to reveal the implicit regularization effect of GF for deep diagonal linear networks. However, the explicit form of the corresponding entropy function for the induced mirror flow dynamics is completely absent. There is even no suggestion about possible properties that the entropy function should have.
In addition, the derivation of the mirror flow form is a direct application of results in Li et al., 2022, and the first point of Proposition 1 can be a direct application of the Euler’s theorem for homogeneous function. Thus I think the technical contributions of this paper are rather limited.
2. As a comparison, Yun et al., 2021 already explicitly characterized the implicit bias of deep diagonal linear networks by using the tensor network formulation developed in their paper.
Specifically, they established the optimization problem with an explicit form of the entropy function that the GF dynamics of deep diagonal linear networks (note that they did not require the parameterization $u^{\odot L} - v^{\odot L}$) aims to solve. They also established the convergence of the dynamics. The only possible weakness of their result is the additional requirement of the initialization, which the authors in this paper are able to relax at the cost of the characterization for explicit form of entropy function. But I cannot view such relaxation as a significant theoretical contribution that is sufficient for this paper to be published in its current version.
**Reference**
Yun et al., 2022. A Unifying View on Implicit Bias in Training Linear Neural Networks.
Questions
1. What are the technical and theoretical contributions of this paper compared to Yun et al., 2021? For example, are results in this paper more general than those in Yun et al., 2021?
2. Can you derive the explicit form of the entropy function of the induced mirror flow dynamics?