Minimax Optimal Rate for Parameter Estimation in Multivariate Deviated Models

We study the maximum likelihood estimation (MLE) in the multivariate deviated model where the data are generated from the density function $(1-\lambda^{\ast})h_{0}(x)+\lambda^{\ast}f(x|\mu^{\ast}, \Sigma^{\ast})$ in which $h_{0}$ is a known function, $\lambda^{\ast} \in [0,1]$ and $(\mu^{\ast}, \Sigma^{\ast})$ are unknown parameters to estimate. The main challenges in deriving the convergence rate of the MLE mainly come from two issues: (1) The interaction between the function $h_{0}$ and the density function $f$; (2) The deviated proportion $\lambda^{\ast}$ can go to the extreme points of $[0,1]$ as the sample size tends to infinity. To address these challenges, we develop the \emph{distinguishability condition} to capture the linear independent relation between the function $h_{0}$ and the density function $f$. We then provide comprehensive convergence rates of the MLE via the vanishing rate of $\lambda^{\ast}$ to zero as well as the distinguishability of two functions $h_{0}$ and $f$.

Paper

Similar papers

Peer review

Reviewer HL4W6/10 · confidence 2/52023-07-02

Summary

The paper studies the optimal rate for multivariate deviated models. Specifically, they consider the model $(1-\lambda) h(x) + \lambda f(x|\mu,\Sigma)$, where $h$ is known and the goal is to estimate the other parameters. The authors propose to use the notion of *distinguishability* and study the convergence rate of parameters using MLE under both distinguishable and non-distinguishable cases. The authors present three pairs of upper and lower bounds for distinguishable, non-distinguishable but $f$ is strongly-identifiable, and distinguishable but $f$ is a family of location-covariance multivariate Gaussian distributions. Experiments are also provided to corroborate their theoretical results.

Strengths

- The paper is clear. The notations are very consistent for such a lengthy work. - There are plenty of explanations and discussions around each main result, making it an interesting paper to read. - To the best of my knowledge, the proofs are technically sound and highly sophisticated. The notion of distinguishability and identifiability seems very suitable and intuitive. - Lower bounds are also presented and match their convergence rate for all three cases considered.

Weaknesses

- The idea of distinguishability is not wholly novel. Similar notions have been used in [1] for a different model, which in turn is derived from the notion of *identifiability* adopted in [2] and many other previous works. - The section for related works is very concise. - Given the classical parametric setting and the MLE estimation, I am not very sure if this paper would be a good match for neurips rather than a more statistically-focused journal. [1] Do, Dat, Nhat Ho, and XuanLong Nguyen. "Beyond black box densities: Parameter learning for the deviated components." Advances in Neural Information Processing Systems 35 (2022): 28167-28178. [2] Nguyen, XuanLong. "Convergence of latent mixing measures in finite and infinite mixture models." (2013): 370-400.

Questions

- I totally understand the logic behind the organization of the sections but I would suggest shortening Section 3 and moving the results in Appendix A to the main content. Theorem 3.3/3.5/3.6 seem to be intermediate results for proving the convergence rates with artificial distances $\mathcal{K}$, $\mathcal{D}$, and $\mathcal{G}$. So I don't quite understand why they should take up nearly three pages, forcing the main results to be postponed to the appendix. - Say we would like to adopt the model to fit some data in practice. How should we obtain $h_0$ which is claimed to be known in this paper? Or $h_0$ could just be good enough and then this model may take over? Would you please comment on how the results may provide possible guidance to the methodology?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

3 good

Contribution

3 good

Limitations

N/A

Reviewer UU3o7/10 · confidence 3/52023-07-05

Summary

In this paper, the authors establish the rate for estimating true parameters in the multivariate deviated model by using the MLE method. They mainly try to address two challenges encountered in deriving the rate of convergence for MLE estimators, i.e. 1) the interaction between the null hypothesis density $h_0$ and the alternative density function $f$, 2) the likelihood of the deviated proportion $\lambda$ vanishing to either extreme points of the interval [0, 1]. To this end, they develop the distinguishability condition to capture the linear independent relation between the function $h_0$ and the density function $f$, and derived the optimal convergence rate of the MLE under both distinguishable and non-distinguishable settings.

Strengths

The paper is well-structured and effectively presents the problem setup, theoretical framework, and main results. The definitions and explanations of key concepts are well presented. The authors address a fundamental statistical problem and provide insights into the behavior of the MLE in the multivariate deviated model. The derived convergence rates and minimax rates contribute to the understanding of parameter estimation and hypothesis testing in complex data scenarios.

Weaknesses

I feel that in the paper the comparison to existing literature is a bit limited. Particularly, how does this paper compare with the current literature on heterogeneous mixture detection? In the experiment section, it seems that the setup is a bit limited with f being Gaussian. It would be great if the authors could show some more numerical results with more expanded scenarios.

Questions

How does this paper compare with the current literature on heterogeneous mixture detection in terms of assumptions and results? Can the results in Section 3.2.2 be extended to non-Gaussian distributions for f?

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

No potential negative societal impact

Reviewer bb5i6/10 · confidence 2/52023-07-06

Summary

## Summary The authors study the minimax rate for parameter recovery in deviated multivariate models. In this setting, we observe samples from a mixture (of *unknown* weight \lambda) of a "null" distribution h_0 and a distribution from a parametric family f( | \mu, \Sigma). The goal is to recover from n samples the (\lambda, \mu, \Sigma). ## Contribution The authors study the MLE performance in various regimes (depending on how "far" is h_0 from the parametric family of distributions. Their analysis is very tight, leading to obtaining the minimax rates.

Strengths

I like the result: it is mathematically clean and leads to tight results. It is always nice to see the minimax rates for new problems.

Weaknesses

I am unfortunately failing to conclude the position of the paper in the literature. Are the authors the first to obtain results in this setting? If yes, please explain (much) more why studying the model is interesting/important. If not, please compare thoroughly the results with the existing ones. Also, are the techniques used new in any way? Or are the results simply an application of known techniques? All in all, I think the paper needs a much more thorough literature review.

Questions

See above.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

2 fair

Contribution

3 good

Limitations

See above.

Reviewer 9DGR6/10 · confidence 4/52023-07-06

Summary

The paper studies the problem of parameter recovery in the multivariate deviated model where the data is generated according to the following distribution: $$ (1 - \lambda) h_0 (x) + \lambda f(x | \mu, \sigma) $$ where $f$ belongs to a mean-variance family and $h_0$ is known. One prominent example of such a family is the family of Gaussian distributions. The paper studies the recovery problem under three settings. The first is the distinguishable setting where $h_0$ and the density $f$ are distinguishable (essentially $h_0$ cannot be written as a linear combination of two distributions from the family the family) where statistical recovery is guaranteed at a $\sqrt{n}$ rate. In the second setting of strong identifiability, $f$ is assumed to belong to the mean-variance family and here the convergence rates depend on the closeness of $\mu_0, \Sigma_0$, the parameters of $h_0$, to $\mu, \Sigma$, the parameters of the unknown mixture component. Finally, in the third setting where the mean-variance family is the Gaussian distribution where different convergence behavior is observed from the strongly identifiable setting where a second order PDE guarantees improved performance. From a technical standpoint, the algorithm (MLE) is analyzed in the following way. First, they show that under some mild assumptions on the function class, the MLE solutions approximate the distribution in Hellinger-distance. Then, by noting the relationship between the Hellinger and TV distances, the paper then shows that the for this class of distributions, the distance between the parameters is upper bounded by a constant multiple of the TV distance. The first step follows by standard empirical process theory. The second, however, relies on some intricate recently developed machinery. Roughly speaking, one first shows that the TV and parameter-distances approximate each other locally where the limit of the neighborhood is taken to $0$. Subsequently, a short analytic argument leads to a global approximation guarantee. This technique, while inspired prior work, still takes significant care to execute. Overall, the results in the paper are interesting and the technical contributions seem strong. The fact that the statistical performance in this setting may be distinguished from algorithms that operate on mixture models where both components are unknown is also intriguing. However, Theorem 3.6 is only proved for the univariate setting (Appendix C3) while the rest of the paper focuses on the multivariate setting.

Strengths

See main review

Weaknesses

See main review

Questions

See main review

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

Yes

Reviewer 1EFM3/10 · confidence 3/52023-07-07

Summary

This paper tackles the issue of parameter estimation in the deviated Gaussian mixture of experts problem using the Maximum Likelihood Estimation (MLE) method. The authors propose new distances and analyze the convergence of MLE under distinguishable and non-distinguishable conditions.

Strengths

This paper is in relatively good shape. The results seem to be solid.

Weaknesses

The major weakness is the novelty. This paper basically considers a much simpler case than the paper https://huynm99.github.io/Deviated_MoE.pdf. They consider multiple $k$ while this paper considers a single $k$. The definitions, results, organization, and even the notations are almost the same. For example, the hellinger distance and TV distance (quite strange to me but adopted by both papers interesting), although this paper changes the distance from $D$.

Questions

1. what is the mini-max rate of this problem? how close is the current rate to the min-max rate? 2. What is the difference and technical novelty compared to https://huynm99.github.io/Deviated_MoE.pdf

Rating

3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

3 good

Contribution

2 fair

Limitations

NA

Authorsrebuttal2023-08-12

We thank Reviewer HL4W for your positive evaluation of our paper after the rebuttal and for maintaining your score (6).

Reviewer bb5i2023-08-14

I thank the authors for their response which addressed my concerns. I am maintaining my score (6), but with low confidence, subject to the authors adding all the above literature comments in their paper.

Authorsrebuttal2023-08-14

We thank Reviewer bb5i for your positive evaluation of our paper after the rebuttal and for maintaining your score of weak accept (6).

Reviewer UU3o2023-08-16

Thank you very much for the response. I will maintain my score.

Authorsrebuttal2023-08-16

Response to Reviewer UU3o

We thank Reviewer UU3o for your positive evaluation of our paper after the rebuttal and for maintaining your score (7). Best, The Authors

Authorsrebuttal2023-08-18

Thank you

We thank Reviewer 9DGR for your positive evaluation of our paper after the rebuttal and for maintaining your score of weak accept (6) with high confidence.

Authorsrebuttal2023-08-19

Dear Reviewer 1EFM, We would like to thank you very much for your feedback, and we hope that our response addresses your previous concerns about our paper. However, as the discussion period is expected to conclude in the next few days, please feel free to let us know if you have any further comments on our work. We would be more than happy to address any additional concerns from you. Thank you again for spending time on the paper, we really appreciate that! Best regards, The Authors

Reviewer 1EFM2023-08-20

Thank you

Dear authors, After reading your response, I think we all agree that the setting of your paper and the paper I listed are quite similar, and some of your discovered results are almost the same (as in your response the $O(n^{-1/2})$ convergence rate). I agree that there are some differences in the proof procedure, otherwise, they should be in one paper. Still, most of your derivations follow almost the same procedure as the paper I mentioned. The differences you mention are mostly driven by some mathematical manipulations. I am afraid the impact of this paper is limited. In summary, I would like to remain my previous ratings. Thank you.

Authorsrebuttal2023-08-21

Response to Reviewer 1EFM

Dear Reviewer 1EFM, Thank you for your response. However, we gracefully disagree with your comments due to the following reasons: **(1)** The objectives of our paper and [1] are totally different. While our paper focuses on establishing **uniform rates** for parameter estimation in multivariate deviated models, the paper [1] concentrates on deriving **point-wise rates** for parameter estimation in deviated Gaussian mixture of experts. In particular, we allow true parameters $G_{\ast}=(\lambda^{\ast},\mu^{\ast},\Sigma^{\ast})$ to vary with the sample size $n$. Meanwhile, ground-truth parameters in [1] remain unchanged with respect to $n$. Thus, our derived parameter estimation rates are uniform, but those in [1] are point-wise. Additionally, the convergence behavior of parameter estimation in our paper is not similar to that in [1]. For instance, under the distinguishable settings, [1] claims that the estimation rates for true parameters are of order $\mathcal{O}(n^{-1/2})$. By contrast, our paper points out that the rate for estimating $(\mu^{\ast},\Sigma^{\ast})$ should be lower than $\mathcal{O}(n^{-1/2})$ since it is determined by the convergence rate of $\lambda^{\ast}$ to zero via the following bound: $\lambda^{\ast}||(\widehat{\mu}_{n}-\mu^{\ast},$ $\widehat{\Sigma}_{n}-\Sigma^{\ast})||=\mathcal{O}(n^{-1/2})$. It is clear that **our rates are sophisticated and able to highlight the implicit interactions between the convergence rates of various parameter estimations, which remains missing in [1]**. To achieve the above rates, we have to face several challenging settings in our proofs. For instance, we first need to make sure that two sequence $G_{n}$ and $G_{\ast,n}$ converge to the same limit $\overline{G}$ under the proposed loss functions. Furthermore, there is still a possibility that the last two components of $G_{n}$ or $G_{\ast,n}$ may not converge to those of $\overline{G}$ under the $2$-norm. Thus, it takes us greater effort to consider all these possible scenarios than in [1], where the authors only need to control the convergence of $G_n$ to $G_{\ast}$. Finally, we would like to emphasize that while we present the minimax lower bound results in Section 4 of our paper, such results remain missing in [1]. **(2)** Let us briefly summarize the literature review for deviated models here, and we would like to refer the reviewer to our general response for further details. The general deviated model is given by: $$p^{\ast}(x) = \lambda^{\ast} h_0(x) + (1-\lambda^{\ast}) f^{\ast}(x),$$ where $h_0$ is known and $(\lambda^{\ast}, f^{\ast})$ are to be estimated from data. (2.1) [2] considers this model specifically when $h_0 = N(0, 1)$ and $f^{\ast} = N(\mu^{\ast}, 1)$ are normal distributions. In the setting $\lambda^{\ast} = n^{-\beta}$ where $\beta \in (0, 1/2)$, they prove that no test can reliably detect $\lambda^{\ast} = 0$ against $\lambda^{\ast} \neq 0$ when $\lambda^{\ast} \mu^{\ast} = o(n^{-1/2})$, while the Likelihood Ratio Test can consistently do it when $\lambda^{\ast} \mu^{\ast} \gtrsim n^{-1/2+\epsilon}$ for any $\epsilon > 0$. However, no guarantee for estimation of $\lambda^{\ast}$ and $\mu^{\ast}$ is provided. (2.2) The uniform convergence of estimating $\lambda^{\ast}$ and $\mu^{\ast}$ is then revisited in [3], in the same setting, where it provides minimax rate and uniform convergence rates for both $\lambda^{\ast}$ and $\mu^{\ast}$ under the $l^2$ estimation strategy. They prove the tight convergence rate for $\lambda^{\ast}$ and $\mu^{\ast}$ when $\lambda^{\ast} |\mu^{\ast}| \gtrsim n^{-1/2 + \epsilon}$ and $|\mu^{\ast}| \gtrsim n^{-1/4}$. However, their technique heavily relies on the properties of the location Gaussian family, which might be difficult to generalize to other settings of kernel densities. **(3)** Minor point: The paper [1] that the reviewer mentioned is actually a draft and it has not been published at any official venues (journals, conferences or even arXiv). With the three reasons (1)-(3), we think that it is not fair to use [1] as the main reason to reject our paper. We hope that the reviewer will consider re-evaluating our paper. **References** [1] H. Nguyen, K. Nguyen, N. Ho. On Parameter Estimation in Deviated Gaussian Mixture of Experts. [2] T. Cai. Optimal detection of heterogeneous and heteroscedastic mixtures. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 2011. [3] S. Gadat. Parameter recovery in two-component contamination mixtures: The l2 strategy. In Annales de l’Institut Henri Poincaré, Probabilitéset Statistiques. Institut Henri Poincaré, 2020. [4] Heinrich, Philippe, and Jonas Kahn. 2018. “Strong Identifiability and Optimal Minimax Rates for Finite Mixture Estimation.” The Annals of Statistics 46 (6A): 2844–70.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC