Thank you for your clarification
We sincerely apologize for misunderstanding your initial comments and greatly appreciate your detailed clarification. Your insights have allowed us to better address your concerns, particularly regarding the assumptions (A1)-(A5). Below, we provide a detailed discussion of these assumptions and their role in our theoretical framework. We are also willing to include these points in the final version of our paper if reviewers deem it beneficial.
### **1. Assumptions (A3) and (A4):**
These assumptions concern the algorithm parameters $\sigma_t$ and $m_t$, which are specified by the user. They are satisfied by many standard FM methods.
### **2. Assumptions (A1), (A2), and (A5):**
These assumptions reflect the properties of the target probability, which is unknown in usual settings. While we acknowledge that verifying these assumptions for specific datasets is infeasible, The primary purpose of the current and many other theoretical work is to compare the potential ability of estimators or learning methods by the worst-case analysis over a function class, revealing the dependence on important parameters such as dimensionality and smoothness degree. . These rates provide a comparative understanding of the estimator's performance.
For instance, as discussed in Section 3.1, KDE with a Gaussian kernel achieves a minimax rate of $O(n^{-4/(4+d)})$, while using an optimal kernel can yield a better rate of $O(n^{-2s/(2s+d)})$ for the densities on $[0,1]^d$ with smoothness $s$. This comparison informs practical choices by highlighting the importance of kernel selection based on expected smoothness. Similarly, our work demonstrates that FM methods attain the almost minimax optimal convergence rate, which is comparable to DM methods, offering key theoretical insights into FM’s ability.
### **3. Theoretical Role of (A1)**
We understand your concern regarding the practicality of (A1). While it may not always align with real-world data, it is critical for ensuring smoothness conditions that allow rigorous analysis. Relaxing this assumption is an important direction for future work. Nonetheless, (A1) does not necessarily impose overly unrealistic conditions; for example, functions with smoothness $s$ but not $s+1$ may still be differentiable almost everywhere (e.g., ReLU is non-differentiable only at the origin). This ensures that the function class considered under (A1) is both theoretically rich and practically relevant.
### **4. Besov space**
To address your question regarding Besov spaces, Tong et al. (2023) did not use such spaces because their analysis did not involve convergence rates in large-sample asymptotics. In contrast, our work focuses on deriving minimax rates, which inherently depend on the smoothness of the target density. Besov spaces, despite their complexity, have been widely recognized as effective tools for formalizing smoothness. They have been extensively used in theoretical studies on approximation and estimation accuracy (DeVore et al., 1992; Donoho & Johnstone, 1995), and we adopt this tradition to characterize FM methods rigorously.
### References:
R. A. DeVore, B. Jawerth and B. J. Lucier. (1992) Image compression through wavelet transform coding. *IEEE Transactions on Information Theory,* 38 (2) 719-746.
Donoho, D. L., & Johnstone, I. M. (1995). Adapting to Unknown Smoothness via Wavelet Shrinkage. *Journal of the American Statistical Association*, *90* (432), 1200–1224.