Summary
This paper explores the Matrix Factorization problem, which involves estimating the matrices $X \in \mathbb{R}^{N \times N}$ and $Y \in \mathbb{R}^{N \times M}$ given the noisy matrix $S = \sqrt{\kappa} X Y + W$.
The focus is on the high-dimensional regime, where $N/M \to \alpha$, and the investigation includes bi-rotationally invariant $Y$ and $W$, as well as symmetric rotationally invariant $X$.
The authors examine rotationally invariant estimators, which are estimators that share the same singular vectors as the noisy matrix $S$.
The paper derives rotationally invariant estimators based on oracle knowledge of the target matrices and demonstrates that they are also Bayes optimal. By assuming concentration and utilizing replica methods, the paper derives explicitly computable estimators from the oracle estimators. The empirical performance of these derived estimators is investigated and shown to closely match that of the oracle estimators. This suggests that the estimators derived using non-rigorous methods from statistical physics are indeed optimal.
Strengths
The low-rank matrix factorization problem with finite-rank matrices is now a well-studied topic. Similarly, the low-rank matrix denoising problem with extensive (diverging) ranks has garnered recent interest. However, results on matrix factorization with extensive ranks have been relatively scarce. This paper aims to fill this gap in the literature by providing results for the matrix factorization problem with diverging ranks and under general rotationally-invariant priors.
The paper is well-written overall, and Section 5 provides a concise and easy-to-follow overview of the derivation of the results, which are otherwise quite complex.
Weaknesses
The main results of the paper, which are the explicitly computed Rotationally Invariant Estimators, rely on non-rigorous methods from statistical physics. While the empirical results are compelling, it would be valuable in the future to establish a more solid theoretical foundation for these estimators.
Moreover, the assumptions made on the matrices $X$ and $Y$ may be considered somewhat unnatural. It would be beneficial for the authors to provide additional motivation as to why the findings of this paper could be of interest to the NeurIPS community beyond the specific problem examined here. This could help clarify the broader significance and potential applications of the research.
Questions
Can the authors expand on the relevance of the methodology developed in the paper for the analysis of the weight matrices of neural networks?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Limitations
Theoretical paper with no immediate negative societal impact.