Response to Reviewer 1EFM
Dear Reviewer 1EFM,
Thank you for your response. However, we gracefully disagree with your comments due to the following reasons:
**(1)** The objectives of our paper and [1] are totally different. While our paper focuses on establishing **uniform rates** for parameter estimation in multivariate deviated models, the paper [1] concentrates on deriving **point-wise rates** for parameter estimation in deviated Gaussian mixture of experts. In particular, we allow true parameters $G_{\ast}=(\lambda^{\ast},\mu^{\ast},\Sigma^{\ast})$ to vary with the sample size $n$. Meanwhile, ground-truth parameters in [1] remain unchanged with respect to $n$. Thus, our derived parameter estimation rates are uniform, but those in [1] are point-wise.
Additionally, the convergence behavior of parameter estimation in our paper is not similar to that in [1]. For instance, under the distinguishable settings, [1] claims that the estimation rates for true parameters are of order $\mathcal{O}(n^{-1/2})$. By contrast, our paper points out that the rate for estimating $(\mu^{\ast},\Sigma^{\ast})$ should be lower than $\mathcal{O}(n^{-1/2})$ since it is determined by the convergence rate of $\lambda^{\ast}$ to zero via the following bound:
$\lambda^{\ast}||(\widehat{\mu}_{n}-\mu^{\ast},$
$\widehat{\Sigma}_{n}-\Sigma^{\ast})||=\mathcal{O}(n^{-1/2})$.
It is clear that **our rates are sophisticated and able to highlight the implicit interactions between the convergence rates of various parameter estimations, which remains missing in [1]**. To achieve the above rates, we have to face several challenging settings in our proofs. For instance, we first need to make sure that two sequence $G_{n}$ and $G_{\ast,n}$ converge to the same limit $\overline{G}$ under the proposed loss functions. Furthermore, there is still a possibility that the last two components of $G_{n}$ or $G_{\ast,n}$ may not converge to those of $\overline{G}$ under the $2$-norm. Thus, it takes us greater effort to consider all these possible scenarios than in [1], where the authors only need to control the convergence of $G_n$ to $G_{\ast}$.
Finally, we would like to emphasize that while we present the minimax lower bound results in Section 4 of our paper, such results remain missing in [1].
**(2)** Let us briefly summarize the literature review for deviated models here, and we would like to refer the reviewer to our general response for further details. The general deviated model is given by:
$$p^{\ast}(x) = \lambda^{\ast} h_0(x) + (1-\lambda^{\ast}) f^{\ast}(x),$$
where $h_0$ is known and $(\lambda^{\ast}, f^{\ast})$ are to be estimated from data.
(2.1) [2] considers this model specifically when $h_0 = N(0, 1)$ and $f^{\ast} = N(\mu^{\ast}, 1)$ are normal distributions. In the setting $\lambda^{\ast} = n^{-\beta}$ where $\beta \in (0, 1/2)$, they prove that no test can reliably detect $\lambda^{\ast} = 0$ against $\lambda^{\ast} \neq 0$ when $\lambda^{\ast} \mu^{\ast} = o(n^{-1/2})$, while the Likelihood Ratio Test can consistently do it when $\lambda^{\ast} \mu^{\ast} \gtrsim n^{-1/2+\epsilon}$ for any $\epsilon > 0$. However, no guarantee for estimation of $\lambda^{\ast}$ and $\mu^{\ast}$ is provided.
(2.2) The uniform convergence of estimating $\lambda^{\ast}$ and $\mu^{\ast}$ is then revisited in [3], in the same setting, where it provides minimax rate and uniform convergence rates for both $\lambda^{\ast}$ and $\mu^{\ast}$ under the $l^2$ estimation strategy. They prove the tight convergence rate for $\lambda^{\ast}$ and $\mu^{\ast}$ when $\lambda^{\ast} |\mu^{\ast}| \gtrsim n^{-1/2 + \epsilon}$ and $|\mu^{\ast}| \gtrsim n^{-1/4}$. However, their technique heavily relies on the properties of the location Gaussian family, which might be difficult to generalize to other settings of kernel densities.
**(3)** Minor point: The paper [1] that the reviewer mentioned is actually a draft and it has not been published at any official venues (journals, conferences or even arXiv).
With the three reasons (1)-(3), we think that it is not fair to use [1] as the main reason to reject our paper. We hope that the reviewer will consider re-evaluating our paper.
**References**
[1] H. Nguyen, K. Nguyen, N. Ho. On Parameter Estimation in Deviated Gaussian Mixture of Experts.
[2] T. Cai. Optimal detection of heterogeneous and heteroscedastic mixtures. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 2011.
[3] S. Gadat. Parameter recovery in two-component contamination mixtures: The l2 strategy. In Annales de l’Institut Henri Poincaré, Probabilitéset Statistiques. Institut Henri Poincaré, 2020.
[4] Heinrich, Philippe, and Jonas Kahn. 2018. “Strong Identifiability and Optimal Minimax Rates for Finite Mixture Estimation.” The Annals of Statistics 46 (6A): 2844–70.