Classification of Heavy-tailed Features in High Dimensions: a Superstatistical Approach

We characterise the learning of a mixture of two clouds of data points with generic centroids via empirical risk minimisation in the high dimensional regime, under the assumptions of generic convex loss and convex regularisation. Each cloud of data points is obtained via a double-stochastic process, where the sample is obtained from a Gaussian distribution whose variance is itself a random parameter sampled from a scalar distribution $\varrho$. As a result, our analysis covers a large family of data distributions, including the case of power-law-tailed distributions with no covariance, and allows us to test recent"Gaussian universality"claims. We study the generalisation performance of the obtained estimator, we analyse the role of regularisation, and we analytically characterise the separability transition.

Paper

Similar papers

Peer review

Reviewer 7S1w7/10 · confidence 3/52023-07-03

Summary

The paper is concerned with binary classification, when the data comes from two point clouds that are superposition of Gaussian distribution. This model allows for data distribution with fat tails. The authors analyse the performance of empirical risk minimization in the high-dimensional regime where the number of training samples and the dimension jointly diverge. Using the replica method from statistical physics, they reduce the computation of e.g the training loss / generalization error to the resolution of self-consistent equations. In the third section of the paper, the authors apply their main result on experiments with synthetic data.

Strengths

The paper is clearly written, and the mathematical derivations are easy to follow. The application of the replica method on a mixture of superposition of Gaussian distributions is new, to the best of my knowledge.

Weaknesses

The paper lacks an experiment on real data to showcase situations in which the data model used in this paper (superposition of Gaussian) is more realistic / useful than simply using a mixture of two Gaussians.

Questions

Could the computations be easily extended to : 1) other types of estimators, e.g. Bayesian estimators that sample from a posterior distribution instead of ERM 2) more generic covariance for the Gaussian distributions. For instance instead of $N(\mu, \Delta I_d)$, use $N(\mu, \Delta \Sigma)$ where $\Sigma is a generic covariance matrix.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

n/a

Reviewer X8465/10 · confidence 2/52023-07-07

Summary

This paper investigates the asymptotic behavior of Generalized Linear Models (GLM) when the number of training samples $n$, and the dimension of the feature-space $d$ both go to infinity, but the ratio $n/d$ is fixed to some known bounded value $\alpha$. Moreover, authors assume the training data points are drawn from a mixture of two heavy-tailed distributions, which is different from the usual Gaussian assumption in most of the existing works (the mentioned heavy-tailed distributions are constructed by combining uncountably infinitely many Gaussian distributions). Paper claims to achieve a non-trivial asymptotic characterization of this problem setting, and also try to validate it via a number of experiments on synthetic data. Paper has a number of shortcomings, therefore my current vote is borderline reject. Presentation of the main results needs to be significantly improved, and also I would like to see the comments from other reviewers with more expertise in this particular field to assess the level of technical contribution in this work.

Strengths

- Paper is well-written (at least in most parts), and the literature review part in the introduction section is very informative. - I have not completely checked the proofs, however, I have not noticed any mathematical mistakes. The technical validity of the theoretical part looks fine. I have not checked the experimental parts.

Weaknesses

- All the theoretical derivations are based on asymptotics, while any result in the non-asymptotic case would be far more interesting. - I suggest presenting the results more formally, i.e., in the form of Theorems, Lemmas, and etc. Otherwise, the actual level of technical contribution in this work becomes hard to assess. Right now, there are no theorems inside the manuscript. Also, the explanations from L.130 to L.158 are vague (please see the questions section). - The main motivation behind this work is to assume non-Gaussian distributions (with a possibly infinite covariance) as the components of the mixture model which generates the data. I am concerned with how much this setting would look imporntant and/or interesting to the community. Due to the Gaussian universality principle, the analysis based on the Gaussian assumptions applies (more or less) to all "Gaussian-like" distributions (which covers almost all distributions with bounded moments) as well. Heavy-tailed distributions with power-law tails which do not have a bounded covariance are of course excluded from this list, but how important are they? IMO, authors have not given enough motivation regarding this issue. - The process which is used to generate the above-mentioned heavy-tailed distributions is very specific: superposition of an uncountably many Gaussians, or equivalently assuming that the covariance matrix of the Gaussian itself is a R.V. with an inverse-Gamma distribution. Authors have not discussed the limitation of this process. How general is it? does it include almost all heavy-tailed distributions? - I have not completely checked the proofs in supp. However, the mathematical tools used for deriving the results are not sophisticated. Not using sophisticated math or not relying on existing elegant theorems is fine, as long as an important problem has been solved or an interesting discovery has been made. This again takes us to a previous comment, on the importance level of this problem setting. I am not familiar with this particular line of research, so I have to wait for other reviewers to comment on that.

Questions

L.130 to L.158: Results are not properly presented. I suggest using a formal theorem and a set of lemmas. L.139, Eq (4): What are $\boldsymbol{g}$ and $\boldsymbol{h}$? L.140, Eq (5): What are $h_{\pm}$, $\omega_{\pm}$, $q$ and ... Actually this list can go on. The main theoretical contributions are presented in Eq (8) and Eq (9). However, the vague explanation preceding them, would impose a huge negative impact on the potential reader.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

-

Reviewer 26qY6/10 · confidence 3/52023-07-10

Summary

The paper is focused on the non-Gaussian mixture model and asymptotical investigation of the asymptotic characterization of the statistics of the empirical risk minimization estimator. The paper takes under consideration the models with two clusters and applies their analysis to the convex loss functions and regularizers. The empirical evaluations investigate the theoretical aspects in practice.

Strengths

- The theoretical analysis of generalized linear models is an exciting direction, especially in the contents of high-dimensional data. - The empirical evaluation of synthetic use cases seems to confirm the theoretical investigations.

Weaknesses

- The paper is chaotic and very difficult to follow. It makes the work very difficult to understand. The listed contusions are not defined to the point. They are focused on studying and analyzing, not on practical outcomes. The work is not even summarized (taking into account that there is still some space in the manuscript, it seems to be strange) in the conclusions section, and limitations are not discussed well. Some explanations of crucial symbols are missing in the paper, and the motivations behind some steps are not explained well. - The proposed theoretical investigation is limited to the models with two classes. How can the results scale to multiclass scenarios? - The empirical evaluation is limited only to artificial cases. The problem investigated by the authors is very practical, and it is crucial to provide some empirical evaluation using real datasets. - There are many ways to go beyond the Gaussian distribution. Normalizing flows may be used to model the distributions for each of the considered clusters as an alternative to this approach. It would be beneficial to discuss this issue in the paper and even provide some empirical comparison to the approach.

Questions

Please refer to the remarks from the Weaknesses section.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

1 poor

Contribution

3 good

Limitations

The paper is not organized properly, it is problematic to identify contributions and a strong plot in the work. Empirical evaluation on real cases is missing.

Reviewer K5kC7/10 · confidence 4/52023-07-19

Summary

The paper derive a theory for training and generalization error when classifying a large number of points from a non-Gaussian high-dimensional data distribution. The data model is a double-stochastic process where a parameter is sampled from a scalar distribution and then a sample is taken from a Gaussian distribution with this parameter as variance. A self-consistent mean-field theory is provided for the case where the number of points is large and proportional to the dimensionality and the equations can be numerically solved for logistic and square losses. The theory is applied to data-sets with finite and infinite covariance, to study the role of regularization, and to estimate the separability threshold for such data. This highlighting both cases of “Gaussian universality” where the results coincide with previous “Gaussian” literature and deviation from such universality.

Strengths

* Originality: the tasks and methods are not new, and previous contributions are very well introduced. The originality of the work lies in the successful calculation of the theory for the non-Gaussian case, which is a valuable contribution. Furthermore, the work provides a basis for analyzing when an extrapolation beyond the Gaussian case is justified. * Significance: the paper is important in highlighting where non-Gaussian data may diverge from the Gaussian case discussed in the literature, and as such it opens a venue for future work to use non-Gaussian analysis of real-worlds data, which is an important direction. The conclusions about test error in Gaussian vs non-Gaussian cases is non-trivial (the inversion between figure 1+2 and 3) and as is the finding or optimal finite regularization value for the non-Gaussian case. * Clarity: the paper is in general well-written and can be served as exemplar for providing complicated theoretical results without sacrificing the clarity of the ideas. * Quality: the paper seems technically very sound, with impressive combination of theory and simulations.

Weaknesses

* Clarity: some of the notations and ideas presented are only hastily introduced, with two prominent examples being “Superstatistical Features” (from the title) and “uncountable superposition” (from the abstract, which seem overly complicated. To me, a presentation through “double stochastic” process (as in my summary) is straight-forward and require no extra jargon. Another avenue for improving the understanding of the reader may lie in “Quadratic loss with ridge regularisation” where results are more amendable to interpretation. The authors should have provided more intuition for those results and furthermore point out where does non-Gaussian enters in the self-consistent equations (i.e., what part is shared with the Gaussian case). * Originality: the main part of the work focus on reproducing known results from Gaussian literature and exploring the deviation from them for the non-Gaussian case. In that sense, there is no originality in this work beyond the (impressive) achievement of providing a theory which describe this non-Gaussian case. * Significance: the work would have been more influential if it provided new tools which can be applied to datasets, where a small number of shape parameters is fit to non-Gaussian data and the ability to classify this data can be predicted from theory and then compared to actual classification of the data. Cases where the Gaussian case predicts the behavior for the non-Gaussian case might deserve a fuller theoretical analysis, perhaps through the analysis suggested for “Quadratic loss with ridge regularisation” above.

Questions

* Why bother with the classifier estimator phi? Is there any reasonable choice beyond sign? * Why do you refer to z* as “the matrix”? * Are all the 8 (or 10) order parameters scalars? * Can you clarify the interpretation of the proximal h and g? Do their distribution is a mean-field version of some real-worlds quantity?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

The authors adequately addressed the limitations of their work. Those include the diagonal structure of the covariance matrix (conditional on the value of delta), the use of K=2 which leads to lack of discussion about the mean of the distribution (because they do not affect anything for K=2 beyond their norm), and the resulting theory solvable only numerically.

Reviewer K5kC2023-08-15

Response to Author Rebuttal

I stand by my assessment that this is a good candidate for acceptance. Looking forward to seeing if the other, less positive reviewers reconsider following the improved presentation and additional results.

Reviewer 26qY2023-08-18

Thank you for rebuttal

I would like to thank the authors for the clarification during the rebuttal. I read the paper one more time, as well as other reviews and comments. After clarification, I appreciate the theoretical contribution of this work. I still think that the paper requires some rewriting to make it more accessible to the larger community. Moreover, I think that empirical evaluation of real cases is possible at least as a showcase for theoretical considerations. I decided to raise my score.

Reviewer 7S1w2023-08-19

Re: Rebuttal by Authors

I thank the authors for their detailed response and appreciate the addition of results for the Bayes-optimal estimator. This leads me to increase my score.

Reviewer X8462023-08-19

I would like to thank the author(s) for their detailed response. After reading the rebuttal and also other reviewers' discussions I decide to slightly raise my rating, but reduce my confidence score.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC