Summary
This paper investigates the learning dynamics of single index models when the input data has a structured covariance. It demonstrates that the commonly used spherical gradient flow fails to learn the target direction even when the spike and target directions are identical. However, the paper introduces an appropriate weight normalization technique that overcomes this limitation and successfully recovers the target direction. It further reveals that leveraging the alignment between the covariance structure and the target direction improves the sample complexity compared to isotropic cases and outperforms lower bounds for rotationally-invariant kernel methods. The paper also suggests future directions, such as studying multi-index models, understanding the limitations from a Correlational Statistical Query (CSQ) perspective, and exploring the effects of different initializations in training networks with multiple neurons.
Strengths
${\bf Originality}$: This paper introduces a new investigation into learning single-index models with a spiked covariance structure, expanding upon prior research that primarily focused on isotropic data. This unique focus on the impact of additional structure in the covariance matrix sets it apart from existing works. Furthermore, the introduction of weight normalization techniques and the exploration of their effects in anisotropic scenarios demonstrate originality in addressing the limitations of conventional spherical gradient dynamics.
$\bf Quality$: This paper demonstrates a high level of quality in terms of its theoretical analysis, rigorous mathematical proofs, and well-supported findings. The authors delve into the dynamics of gradient flow, leveraging mathematical techniques to derive insights into the behavior of the learning algorithm. The use of appropriate assumptions and the incorporation of empirical covariance preconditioning showcase the paper's thoughtful methodology and attention to detail.
$\bf Clarity$: The paper is written in a clear and accessible manner, making complex concepts and mathematical formulations understandable to the reader. The abstract, introduction, and conclusion concisely summarize the objectives, methods, and key findings of the study. The organization of the paper and the logical flow of ideas contribute to its overall clarity, enhancing the reader's comprehension of the research.
$\bf Significance$: The paper holds good implications for the understanding of learning dynamics in the presence of structured covariance matrices. By uncovering phenomena where conventional approaches fail, such as the inability of spherical gradient dynamics to recover the target direction, the paper sheds light on the limitations and challenges associated with anisotropic data. Moreover, the proposed weight normalization techniques, insights into improved sample complexity, and the comparison to rotationally invariant kernel methods highlight the practical relevance and potential impact of the findings.
Weaknesses
${\bf (1)}\ \textbf{The training procedure considered in the paper is not practical:}$ The authors focus on a two-step training procedure that deviates from the standard gradient descent (GD) commonly used in practice. It is challenging to assess the extent to which the results obtained from this training procedure translate to impacts on standard training algorithms.
${\bf (2)} \ \textbf{Differences from the existing works:}$ Some existing works, such as Refs [BBSS22] [BAGJ21], have established some results on learning the single-index model using neural networks. It seems that the authors don't clarify the difference of this work from the existing ones.
${\bf (3)}\ \textbf{Practical implications and applications:}$ Although the authors have successfully established strong theoretical results for learning single-index models using two-layer neural networks, the practical implications of this work on real-world applications of deep learning remain somewhat unclear. Further clarification regarding the practical relevance and potential impact of the findings would greatly benefit readers seeking to understand the practical implications of the research in the field of deep learning.
Questions
$\bf Q1.$ What happens if assuming $g$ and $\phi$ to be differentiable? Can you get stronger results?
$\bf Q2.$ Is it possible to extend the current analysis to the multiple-index models? What are the main technical difficulties?
Rating
5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
Not applicable