Multinomial Logistic Regression: Asymptotic Normality on Null Covariates in High-Dimensions

This paper investigates the asymptotic distribution of the maximum-likelihood estimate (MLE) in multinomial logistic models in the high-dimensional regime where dimension and sample size are of the same order. While classical large-sample theory provides asymptotic normality of the MLE under certain conditions, such classical results are expected to fail in high-dimensions as documented for the binary logistic case in the seminal work of Sur and Cand\`es [2019]. We address this issue in classification problems with 3 or more classes, by developing asymptotic normality and asymptotic chi-square results for the multinomial logistic MLE (also known as cross-entropy minimizer) on null covariates. Our theory leads to a new methodology to test the significance of a given feature. Extensive simulation studies on synthetic data corroborate these asymptotic results and confirm the validity of proposed p-values for testing the significance of a given feature.

Paper

References (31)

Scroll for more · 19 remaining

Similar papers

Peer review

Reviewer U9Cr7/10 · confidence 3/52023-07-04

Summary

The paper aims to study the statistics of the multinomial logistic MLE on null covariates in a multiclass classification setting in the asymptotic proportional regime, i.e., in the limit in which the size of the dataset $n$ and of the dimension of the covariates $p$ are both sent to infinity but with finite ratio. The work extends previous results by Sur and Candès in which the normality of the MLE on null covariates was proven in the case of binary classification and introduces a proper statistic for a $\chi^2$ test on the relevance of a feature.

Strengths

The paper is clearly and carefully written, the validity conditions of the results are clearly stated and the outcome of the theoretical analysis is supported by robust numerical evidence. The characterization of the statistics of the MLE on null covariates in the case of multiclass classification in the high dimensional limit is timely. The authors also compare the high-d and the classical theory, showing the remarkable difference in the provided predictions. Finally, they also test the validity of their results beyond the Gaussian hypothesis adopted for the derivation of their theorem.

Weaknesses

A possible theoretical weakness of the paper is the fact that the existence of the MLE is assumed (Assumption 2.4) and not fully characterized. This "weakness" is acknowledged by the authors themselves in the final section of the main text.

Questions

Out of curiosity, the authors state that they expect their results to hold for covariates with distribution having "sufficiently light" tails. Could they comment on this? Do they expect, for example, sub-gaussianity to be sufficient for their results to hold? As a (very) minor remark, Theorem 2.1 and Theorem 2.2 exhibit the quantities $\mathsf y_i-\hat{\mathsf p}_i\equiv -\mathsf g_i$, but a different notation is used. The analogy of the two results might appear more evident at first sight using the same notation in both theorems (eg replacing $-\mathsf y_i+\hat{\mathsf p}_i$ with $\mathsf g_i$ in the first theorem).

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

The work is of theoretical nature and all limitations are within the well specified hypotheses of the reported theorems.

Reviewer uvEK7/10 · confidence 2/52023-07-06

Summary

This paper studies the asymptotic distribution of multinomial logistic regression in the high dimensions. Under the assumption of a Gaussian design (and a few other assumptions), it establishes and characterizes the asymptotic normality of the coefficient of a null feature. This extends previous results for high-dimensional binary logistic regression to $K \geq 3$ number of classes. This result enables proper significant testing in high-dimensions, for which the classical fixed-$p$ asymptotic fails to control the type-I error.

Strengths

1. The paper extends previous results for high-dimensional logistic regression to the multi-class setting. 2. The paper is well-written and the technical content is presented rigorously. 3. The simulations results seem to match the asymptotic theory very well.

Weaknesses

1. The hypothesis $H_0$ seems stronger than the hypothesis $H_0'$: the population coefficient (in terms of KL projection) for feature $j$ is zero. Supposedly $H_0'$ reduces to $H_0$ when the model is well-specified. In terms of the data generating mechanism given by Assumption 2.1, it seems possible that when $y_i = f(U_i, x_i^{T} B^{\ast})$ does not hold, a feature that does not satisfy $H_0$ may still have zero as its population "true" coefficient. I am wondering whether $H_0$ can be suitably weakened. 2. The result relies on a random Gaussian design, although it is suggested that the result probably extends to other random designs with light tails. 3. The paper does not characterize the distribution for non-null features, so the study of parameter inference for high-dimensional multinomial logistic regression is still incomplete.

Questions

1. Per my first comment on the "Weaknesses", can authors comment on the extent to which $H_0$ can be weakened? 2. To what extent is the "universality" expected to hold for this problem? Does the asymptotic normality fail under a random, heavy-tailed design? What about fixed designs?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

I do not foresee any potential negative societal impact of this work. In terms of limitations, I think obtaining results for the non-null features would make this work a much stronger paper.

Reviewer 7JUy7/10 · confidence 4/52023-07-07

Summary

In this paper, the authors provide asymptotic characterization of the behavior of the maximum likelihood estimator (MLE) of multinomial logistic model (with more than two classes), in the high-dimensional regime where the dimension and the sample size of data go to infinity at the same rate. Under some technical assumptions (that may need some further elaborations), this paper develops asymptotic normality and asymptotic chi-square results for the multinomial logistic MLE on null covariates (see Theorem 2.1 and 2.2). The proposed results can be used for statistical inference and significance test of some specific features. Numerical experiments on synthetic data are provided in Section 3 to validate the proposed theory.

Strengths

This papers focuses on the fundamental and important problem of MLE of multinomial logistic model in the modern high-dimensional regime. The proposed theory improves prior art in characterizing the asymptotic normality and asymptotic chi-square results, both of significance to statistics and ML. The paper is in general well written and easy to follow.

Weaknesses

I do not have strong concerns to raise for this paper. See below for some detailed comments and/or questions.

Questions

1. While I am almost fine with Assumption 2.1 and 2.2 (just being curious, is some upper bound on the spectral norm of the covariance $\Sigma$ needed? or it is just a matter of scaling with respect to $n,p$?), I am a bit confused by Assumption 2.3 and 2.4: are they something intrinsic or for the ease of technical analysis? What happens if, say Assumption 2.3 is violated? Can we have something similar but just more involved or the MLE is totally different? Also, Assumption 2.4 is a bit misleading, in the sense that the assumption is not instinct, and should perhaps be reduced into some assumption on the dimension ratio $p/n$ and/or statistics of the data? I believe it makes more sense to assume something like "the dimension ratio $p/n$, covariance $\Sigma$ and xxx satisfy that xxx". I am also confused by the paragraph after Assumption 2.4 and I am not sure the convergence of some multinomial regression solver can be used as a rigorous theoretical indicator. The algorithm may converge (or believed to converge) due to many reasons. Perhaps some better (numerical) criterion can be proposed by, e.g., checking the gradient and/or Hessian of the point of interest. 2. for the sake of presentation and use, it be helpful to present Theorem 2.2 and the estimation of $\Omega_{jj}$ in form of an algorithm. 3. Almost nothing is mentioned for the proof of the theoretical results (Theorem 2.1 and 2.2): is the proof technically challenging or contains some ingredients and/or intermediate results that may be of independent interest? Could the authors elaborate more on this?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

This paper is primarily of a theoretical nature, and I do not see any potential negative societal impact of this work.

Reviewer FsXG6/10 · confidence 3/52023-07-07

Summary

This paper studies the asymptotic distribution of the MLE of the multinomial logistic regression model when the sample size and the number of parameters are of the same order. The validity of the asymptotic theories is evaluated through extensive simulation studies. The paper is overall very well written and the presentation is clear.

Strengths

The paper is very well written, and the problem under consideration is of great interest. Theoretical results are sound and the numerical experiments are sufficient.

Weaknesses

One of the motivating example of the paper is to study classification with 3 or more classes. And one of the most important goal in classification is the classification error or AUC. In my opinion, the impact of the proposed method on classification errors should be evaluated through simulation studies and some real data examples.

Questions

1. Related to my question in the weakness section, one could use the proposed tests to exclude some noise covariates in the multinomial regression and see how much improvements can be achieved in terms of classification errors. Preferably, one should compare its performance to some existing methods. 2. I would strongly suggests applying the proposed methodology to some real-world benchmark data set. The statistical tests may provide additional insights and the classification errors can be evaluated through cross-validation. 3. The paper's theoretical framework appears to heavily rely on the seminal work of Sur and Candes (2019), because of which I still have some reservation on the novelty of the theory. To address concerns about the novelty of the theory, it would be wise to explicitly highlight the unique theoretical challenges posed by the multinomial logistic regression model compared to the more commonly studied binary logistic regression model.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

NA

Reviewer 7JUy2023-08-14

I thank the authors for the rebuttal and detailed clarifications. And I maintain my positive rating on this paper.

Reviewer uvEK2023-08-14

I appreciate the reply from the authors. I feel that my concerns have been properly addressed. I am raising my score to "Accept".

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC