Summary
In this paper, the authors provide asymptotic characterization of the behavior of the maximum likelihood estimator (MLE) of multinomial logistic model (with more than two classes), in the high-dimensional regime where the dimension and the sample size of data go to infinity at the same rate.
Under some technical assumptions (that may need some further elaborations), this paper develops asymptotic normality and asymptotic chi-square results for the multinomial logistic MLE on null covariates (see Theorem 2.1 and 2.2). The proposed results can be used for statistical inference and significance test of some specific features. Numerical experiments on synthetic data are provided in Section 3 to validate the proposed theory.
Strengths
This papers focuses on the fundamental and important problem of MLE of multinomial logistic model in the modern high-dimensional regime.
The proposed theory improves prior art in characterizing the asymptotic normality and asymptotic chi-square results, both of significance to statistics and ML.
The paper is in general well written and easy to follow.
Weaknesses
I do not have strong concerns to raise for this paper.
See below for some detailed comments and/or questions.
Questions
1. While I am almost fine with Assumption 2.1 and 2.2 (just being curious, is some upper bound on the spectral norm of the covariance $\Sigma$ needed? or it is just a matter of scaling with respect to $n,p$?), I am a bit confused by Assumption 2.3 and 2.4: are they something intrinsic or for the ease of technical analysis? What happens if, say Assumption 2.3 is violated? Can we have something similar but just more involved or the MLE is totally different? Also, Assumption 2.4 is a bit misleading, in the sense that the assumption is not instinct, and should perhaps be reduced into some assumption on the dimension ratio $p/n$ and/or statistics of the data? I believe it makes more sense to assume something like "the dimension ratio $p/n$, covariance $\Sigma$ and xxx satisfy that xxx". I am also confused by the paragraph after Assumption 2.4 and I am not sure the convergence of some multinomial regression solver can be used as a rigorous theoretical indicator. The algorithm may converge (or believed to converge) due to many reasons. Perhaps some better (numerical) criterion can be proposed by, e.g., checking the gradient and/or Hessian of the point of interest.
2. for the sake of presentation and use, it be helpful to present Theorem 2.2 and the estimation of $\Omega_{jj}$ in form of an algorithm.
3. Almost nothing is mentioned for the proof of the theoretical results (Theorem 2.1 and 2.2): is the proof technically challenging or contains some ingredients and/or intermediate results that may be of independent interest? Could the authors elaborate more on this?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
This paper is primarily of a theoretical nature, and I do not see any potential negative societal impact of this work.