Summary
The authors present an investigative study into the learning dynamics of neural networks, specifically low-rank neural networks. The motivation is that training dynamics of neural networks are not fully understood. Low-rank models have practical advantages (time / memory). This work provides a theoretical and empirical overview of the learning dynamics of neural networks with a focus on low-rank models.
Strengths
The main strength of this paper is its significance, i.e. the attempt to explain learning dynamics of neural networks from both a theoretical and empirical perspective. This is a worthy and valuable endeavour which will be of great interest to the field if done correctly and thoroughly. Unfortunately, there are also significant weaknesses, as we will see, which weight heavily against the strengths.
Weaknesses
There are a number of critical weaknesses present in this paper that severely reduce its value. Perhaps there is a valuable message in there but the presentation, clarity and writing make it impossible to recommend this as an impactful paper. Here are some of the weaknesses detailed:
- Flow and writing can be greatly improved. Very abrupt changes between sections (e.g. from 1 to 2). The sections do not have self-contained introductions to help readers orient themselves and understand the reasons and motivations for choices made.
- Related to the first point, formulae are presented without sufficient discussion. Many sections (notably 2 and 3) read like a series of statements rather than a well-constructed and clearly motivated argument. I encourage the authors to improve the flow to help the reader understand the context
- Formatting can be improved. What is *error*? Overview of results in lines 77-88 can be greatly improved with respect to formatting.
- Missing definition, what is BPTT in line 104?
- Lack of discussion on the connection between the presented theoretical and empirical results.
- Stated contributions are not found in the paper (effect of stride is in supplementary material and I could not find sequence truncation experiments anywhere).
Questions
- Why are only leaky-ReLUs investigated, what is the motivation? And how does this differ from standard ReLUs?
- What is the theoretical connection to bottleneck layers bounding the entire network? Why is this limited to bottleneck layers and not the narrowest part of the network?
- What is the impact of training data on the learnt rank? Does the implicit dimension of the dataset influence the rank? With reference to [1]
- Why is the PDF formatted as an image and not as a standard NeurIPS PDF output? The links to section labels and reference links do not work.
[1] Li, C., Farkhoor, H., Liu, R., & Yosinski, J. (2018). Measuring the intrinsic dimension of objective landscapes. arXiv preprint arXiv:1804.08838.
Rating
3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
The authors discuss only one limitation of their work: the relationship between tasks and low-rank emergence. However, there are a number of limitations that need to be discussed. Not only have the experiments scratched the surface of the possible signals from learning dynamics we can collect and paint us a limited picture of what is happening but the results themselves also make us pose new questions that remain undiscussed. Refer to my questions to the authors for some possible discussion points that would improve this paper greatly and help place it within the existing literature and highlight the most obvious next steps to build on this work for future studies.