Summary
This paper investigates the asymptotic behavior of Generalized Linear Models (GLM) when the number of training samples $n$, and the dimension of the feature-space $d$ both go to infinity, but the ratio $n/d$ is fixed to some known bounded value $\alpha$. Moreover, authors assume the training data points are drawn from a mixture of two heavy-tailed distributions, which is different from the usual Gaussian assumption in most of the existing works (the mentioned heavy-tailed distributions are constructed by combining uncountably infinitely many Gaussian distributions).
Paper claims to achieve a non-trivial asymptotic characterization of this problem setting, and also try to validate it via a number of experiments on synthetic data. Paper has a number of shortcomings, therefore my current vote is borderline reject. Presentation of the main results needs to be significantly improved, and also I would like to see the comments from other reviewers with more expertise in this particular field to assess the level of technical contribution in this work.
Strengths
- Paper is well-written (at least in most parts), and the literature review part in the introduction section is very informative.
- I have not completely checked the proofs, however, I have not noticed any mathematical mistakes. The technical validity of the theoretical part looks fine. I have not checked the experimental parts.
Weaknesses
- All the theoretical derivations are based on asymptotics, while any result in the non-asymptotic case would be far more interesting.
- I suggest presenting the results more formally, i.e., in the form of Theorems, Lemmas, and etc. Otherwise, the actual level of technical contribution in this work becomes hard to assess. Right now, there are no theorems inside the manuscript. Also, the explanations from L.130 to L.158 are vague (please see the questions section).
- The main motivation behind this work is to assume non-Gaussian distributions (with a possibly infinite covariance) as the components of the mixture model which generates the data. I am concerned with how much this setting would look imporntant and/or interesting to the community. Due to the Gaussian universality principle, the analysis based on the Gaussian assumptions applies (more or less) to all "Gaussian-like" distributions (which covers almost all distributions with bounded moments) as well. Heavy-tailed distributions with power-law tails which do not have a bounded covariance are of course excluded from this list, but how important are they? IMO, authors have not given enough motivation regarding this issue.
- The process which is used to generate the above-mentioned heavy-tailed distributions is very specific: superposition of an uncountably many Gaussians, or equivalently assuming that the covariance matrix of the Gaussian itself is a R.V. with an inverse-Gamma distribution. Authors have not discussed the limitation of this process. How general is it? does it include almost all heavy-tailed distributions?
- I have not completely checked the proofs in supp. However, the mathematical tools used for deriving the results are not sophisticated. Not using sophisticated math or not relying on existing elegant theorems is fine, as long as an important problem has been solved or an interesting discovery has been made. This again takes us to a previous comment, on the importance level of this problem setting. I am not familiar with this particular line of research, so I have to wait for other reviewers to comment on that.
Questions
L.130 to L.158: Results are not properly presented. I suggest using a formal theorem and a set of lemmas.
L.139, Eq (4): What are $\boldsymbol{g}$ and $\boldsymbol{h}$?
L.140, Eq (5): What are $h_{\pm}$, $\omega_{\pm}$, $q$ and ... Actually this list can go on.
The main theoretical contributions are presented in Eq (8) and Eq (9). However, the vague explanation preceding them, would impose a huge negative impact on the potential reader.
Rating
5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.
Confidence
2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.