Bottleneck Structure in Learned Features: Low-Dimension vs Regularity Tradeoff

Previous work has shown that DNNs with large depth $L$ and $L_{2}$-regularization are biased towards learning low-dimensional representations of the inputs, which can be interpreted as minimizing a notion of rank $R^{(0)}(f)$ of the learned function $f$, conjectured to be the Bottleneck rank. We compute finite depth corrections to this result, revealing a measure $R^{(1)}$ of regularity which bounds the pseudo-determinant of the Jacobian $\left|Jf(x)\right|_{+}$ and is subadditive under composition and addition. This formalizes a balance between learning low-dimensional representations and minimizing complexity/irregularity in the feature maps, allowing the network to learn the `right' inner dimension. Finally, we prove the conjectured bottleneck structure in the learned features as $L\to\infty$: for large depths, almost all hidden representations are approximately $R^{(0)}(f)$-dimensional, and almost all weight matrices $W_{\ell}$ have $R^{(0)}(f)$ singular values close to 1 while the others are $O(L^{-\frac{1}{2}})$. Interestingly, the use of large learning rates is required to guarantee an order $O(L)$ NTK which in turns guarantees infinite depth convergence of the representations of almost all layers.

Paper

References (27)

Scroll for more · 15 remaining

Similar papers

Peer review

Reviewer expT5/10 · confidence 4/52023-07-02

Summary

This work studies the leading order expansions in $L$ of the representation cost for a large $L$, where $L$ is the depth of the model. The work concludes that there is a low-dimension and regularity tradeoff as one varies the depth

Strengths

I think the main strength of the work is the novelty of the message it delivers: there is a low-dimension and regularity tradeoff as one varies the depth. The simplicity of the proofs is a plus. Another important contribution is the technique to expand representation loss for a large depth, though I believe to require a little more technical soundness. See the weakness section Another interesting insight of the paper is the concept of "symmetry learning," and relating the "spurious symmetries" to the representation cost is also insightful At this moment, I have some reservation about accepting this paper. If the authors answer my questions below in a satisfactory manner, I would be happy to accept it

Weaknesses

There are quite a few aspects that I think this work can be improved. 1. Let $R(L)$ denote the representation cost as a function of $L$, the depth. The first step of the theory is to expand $R$ in $1/L$, which feels unjustified, why is this function Taylor-expandable? Why can we treat $L$ as a continuous variable? These points are not crucial problems in my opinion, but the authors do not explain these with sufficient emphasis, given how important they are for the results 2. In my opinion, the most problematic aspect is the fact that it studies the representation cost of an almost infinite depth neural network at a finite weight decay (this also relates to point 1). See https://arxiv.org/abs/2202.04777. This reference (for example, see proposition 3 / theorem 3) implies that for a sufficiently deep neural network and a fixed weight decay, the global minimum is the origin ($\theta=0$), independent of the depth $L$. As one decreases the depth, the global minimum jumps from the origin to a nonzero value. This means that if we restrict to the global minimum of the cost function $C$, the representation cost is not a differentiable function in $L$, and thus it is not Taylor-expandable. I think the authors need to explain why the expansion remains reasonable/correct in light of this result 3. Given how strong the assumptions are (point 1), I think the authors really should have carried out some experiments to directly check the prediction. In particular, I believe the authors should present a numerical example where the author directly compute the first and second order terms of $R$ in $1/L$ and show how they change for real neural networks to make the arguments of the paper more convincing.

Questions

See weakness

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

2 fair

Presentation

4 excellent

Contribution

3 good

Limitations

See weakness

Reviewer tanV6/10 · confidence 2/52023-07-07

Summary

This paper looks at the representation cost $R(f)$ of networks as the $L_2$ norm of all of the parameters. The paper claims that for deep neural networks, computing the minimum cost representation of a function is not tractable. Hence, the look at the representation cost normalized by the number of layers $L$ as $L \to \infty$. The paper then computes the Taylor expansion with respect to $1/L$ at $L = \infty$. Using this expansion, they recover three terms, $R^{(0)}(f),R^{(1)}(f),R^{(2)}(f)$. They note that $R^{(0)}(f)$ is used as a regularize before and corresponded to regularizing this notion of rank called the bottleneck rank. However, they argue that this is not enough and present results that determine the type of regularity imposed by using $R^{(1)}(f), R^{(2)}(f)$ as regularizers as well.

Strengths

The paper is very interesting, as understanding the representation of functions via neural networks is very important. The description of the two new terms $R^{(1)}(f), R^{(2)}(f)$ are explored in depth from a variety of different angles. Hence I think it makes a good contribution to the understanding of neural representation learning.

Weaknesses

My main issue with the paper is its presentation. The paper feels very disjointed. It would be helpful if the authors included more connecting discussion and discussed the high level picture in more detail.

Questions

1. On line 91, what is the infinite width representation cost? Typos 1. Line 156 instead of "of he task" -> "of the task"

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

N/A

Reviewer SexW8/10 · confidence 2/52023-07-08

Summary

This paper introduces two corrections to previous work that relates the representation cost of infinite-depth nonlinear networks to the bottleneck rank. These two corrections are obtained by Taylor expansion of the representation cost. In addition to the previously-known rank term $R^{(0)}(f)$ that is directly related to the bottleneck rank, the authors consider two corrections (1) $R^{(1)}(f)$, which bounds the pseudo-determinant of the Jacobian, and (2) $R^{(2)}(f)$, which plays a role in the intermediate representations in the hidden layer of the network. The authors specifically discuss several theoretical properties of $R^{(1)}(f)$. This term acts as a notion of regularity in the learning process of deep nonlinear networks, and it explains various interesting properties, such as how deep networks trained with conventional L2-regularized loss are uniquely determined (i.e., which $f$ among candidates with the same $R^{(0)}(f)$ is selected) and why the rank of the true function is not underestimated. The authors also briefly demonstrate the relationship between the curvature of the domain manifold and $R^{(1)}$ using the rank of the identity function, and discuss the connection between the learning rate and the Lipschitzness of learned deep networks with the results of the neural tangent kernel. Finally, the authors discuss the dynamics of hidden representations, which in the few initial layers the representations undergo changes to minimize $R^{(0)}$, then, as they move into the low-dimensional space, they change smoothly influenced by $R^{(1)}$ and $R^{(2)}$, i.e., exhibit the bottleneck structure of networks.

Strengths

This paper is a direct follow-up to [1], and since I am not very familiar with [1], my evaluation may not be very accurate. Theoretical understanding of deep nonlinear networks is a highly significant field and holds considerable interest among the NeurIPS audience. For me, the theoretical results of this paper are highly novel and intriguing. The paper successfully modifies the results of [1], and provides a rigorous yet intuitive explanation of why neural networks trained with our L2-regularized loss do not underestimate the rank of the true function. Additionally, it demonstrates that under some conditions, the representation dynamics of deep nonlinear dynamics exhibit a bottleneck structure, which consists of a sequence of long bottleneck dimensional representations. This result is also noteworthy. *** [1] Arthur Jacot. Implicit bias of large depth networks: a notion of rank for nonlinear functions. In The Eleventh International Conference on Learning Representations, 2023.

Weaknesses

I couldn't find any major drawbacks in this paper. The frequent mention of $R^{(2)}$ in Remark 10, before its properties were clearly stated, was somewhat confusing to me, although it is not a critical issue. There are just a few very minor typos: In line 32, “… not control the the …” In line 34, “… formalizes the a …” In line 118, “… because the the …” In line 156, “… $k^{*}$ of he task ..."

Questions

I couldn't find any major drawbacks in this paper. One small question is: I would like to hear the authors’ thoughts on whether the optimal latent dimensionality of a well-trained auto-encoder $f$, which is likely to be similar to the identity $id$, might be associated with $R^{(0)}(f;\Omega)$ and $R^{(1)}(f;\Omega)$, for the case when $Rank_{J}(id;\Omega) = Rank_{BN}(id;\Omega)$, and if possible, when $Rank_{J}(id;\Omega) < Rank_{BN}(id;\Omega)$

Rating

8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

4 excellent

Limitations

The authors do not explicitly address the limitations. However, in some of the theorems/propositions, they briefly mention the limitations of the proposed one, indicating what it can only explain or what constraints it has.

Reviewer dQhS6/10 · confidence 3/52023-07-20

Summary

This paper extends the theoretical framework around the representation cost of DNNs as defined in previous work. The authors introduce two corrections to the existing infinite depth description, namely two regularity measures $R^{(1)}$ and $R^{(2)}$, which balance against the dominating rank bias $R^{(0)}$. The authors propose that these regularity measures prevent networks from underestimating the 'true' bottleneck rank (BN-rank) and argue that large learning rates can also induce a bias toward regularity. The paper provides a theoretical description of the limiting representation geodesics as depth approaches infinity, proving under certain conditions a bottleneck structure in the learned representations. Simple experiment is presented to test the network's ability to learn underlying symmetries.

Strengths

* The paper makes a great analysis of the representation cost under the previous theoretical framework, better explaining some observed results and conjecture. * The theoretical arguments are well supported, with clear intuition and explanation. * Different aspects are considered in the paper such as the effect of learning rate and the representation geodesics which is particularly interesting.

Weaknesses

* The definition of the term "regularity" requires a more specific definition for the problem. It is claimed in line-32 that "this notion of rank does not control the regularity of f". Intuitively, the rank $R^{(0)}$ still measures some notion of complexity of the function as a higher-rank matrix can represent a more complex linear transformation. The corrections are other levels of complexity measures. I guess the "regularity" is more like how simple a function is with the same rank? * Theorem 9 proves bottleneck structure for convergent function sequence with an infinite depth limit. But in practice, the network is trained with finite data points and depth. Could the authors comment on how the conclusions would change for practical networks? * Analyze of the learning rate appears somewhat disjointed from the rest of the paper. Could there be a more integrated discussion about its role? * I found the numerical experiment a bit hard to follow and would appreciate it if more explanation on the idea could be added. Additionally, is it applicable for some experiments on justifying the theoretical findings, such as the effect of learning rate? * Minor typos: * Line32: "...control the the regularity..." * Line34: "This formalizes the a balance..." * Line156: "...$k^*$ of he task..."

Questions

* Can these findings be used to improve the training of DNNs, for example, by adjusting the learning rate or regularization? * Considering that the argument in Proposition 7 is incomplete, how much further work would be needed to provide a complete argument?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

Included in weakness.

Area Chair tJQJ2023-08-11

Hi all, Thanks for serving as the reviewers for this submission. As the authors have already turned in their responses. It is our turn to start the further discussion. Here is a to-do list: (1) Please acknowledge the authors when you finish reading their responses. (2) Please indicate whether you have any further questions for the authors such that they can continue to response. (3) Please indicate whether you are willing to change the ratings. Best AC

Reviewer expT2023-08-18

reply

Thanks for the detailed response. I am partially satisfied with the response, and so I raise the score to 5. What I believe the author should have done much more and better is to include more numerical results to validate the essential predictions of the theory and illustrate its significance

Reviewer dQhS2023-08-21

Thanks for the author's response and they have addressed most of my concerns. I think it would be better to include more experiments to justify the theoretical findings.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC