Dear Area Chair,
Thank you for your insightful comments and careful reading of our submission. We’d like to address your questions as follows:
1. We agree with your observation regarding the need for a discussion on the original sources for results similar to Propositions 2.1 and 3.1, and we will incorporate this in the next version of the paper. Specifically, for Proposition 3.1, we will move the citation of [1], currently in the appendix, to the main body as a standard reference on DP weak convergence to the true generating distribution. Regarding Proposition 2.1, we acknowledge that pinpointing the exact origin of the classical interpretation of Ridge regression as Bayesian linear regression with a *parametric* standard normal prior on the coefficients is challenging. However, as noted after Proposition 2.1, our result offers a novel Bayesian interpretation: by taking an optimization-centric view and placing a *nonparametric* prior directly on the data-generating distributions (instead of on the coefficients, which we optimize), we also arrive at Ridge regularization. We believe this result highlights the fundamental connections between optimization, decision theory, and Bayesian inference, and we will ensure this is clearly explained in the revised paper.
2. We also agree that our paper currently lacks an explicit discussion of our reasoning for relating $\hat V$ to $V_{\boldsymbol\xi^n}$ instead of directly to $\phi(\mathcal R_{p_\star}(\cdot))$. We will address this in the next version of the paper with the following clarifications. As you noted, one's ultimate goal may be to ensure the convergence of $\hat V$ to $\phi(\mathcal R_{p_\star}(\cdot))$, which involves three layers of approximation: from the finite sample size $n$, from the random measure truncation threshold $T$, and from the number of MC samples $N$. In Section 3, we address the first layer by studying the convergence of $V_{\boldsymbol\xi^n}$ to $\phi(\mathcal R_{p_\star}(\cdot))$. In Section 4, we focus on the latter two layers, *given a fixed sample size approximation determined* by $n$, by studying the convergence of $\hat V$ to $V_{\boldsymbol\xi^n}$. Thus, within this logical chain, $V_{\boldsymbol\xi^n}$ serves as a bridging quantity between $\hat V$ and $\phi(\mathcal R_{p_\star}(\cdot))$. The results from Sections 3 and 4 can then be combined to directly establish the convergence of $\hat V$ to $\phi(\mathcal R_{p_\star}(\cdot))$, which involves choosing $T$ and $N$ as functions of $n$ to ensure that the right-hand side of the first equation in the statement of Lemma 4.3 converges to zero. We will emphasize this crucial point in our next revision.
Thank you again for your valuable feedback,
The Authors
*[1] S. Ghosal and A. Van der Vaart. Fundamentals of nonparametric Bayesian inference, volume 44. 378 Cambridge University Press, 2017.*