Rebuttal by Authors
Thank you for your constructive and positive comment with an important question.
__Q1 Is the NPL posterior robust to model misspecification compared to the regular Bayesian posterior on $\theta$?__
Thank you to the reviewer for posing an important and insightful question. You are indeed correct that the parametric posterior and NPL posterior target the same parameter, that is $\theta^*$ which minimizes the KL divergence between the true $F^*$ and the model $F_\theta$ [Walker (2013)].
However, we highlight that NPL elicits a nonparametric prior, which induces a
posterior distributions on $\theta^*$ that has superior asymptotic properties to the regular Bayesian posterior when the model is misspecified. Intuitively, the reason for this subpar performance for parametric Bayes is that the posterior is computed assuming there exists $\theta^*$ such that $F_{\theta^*} = F^*$. On the other hand, NPL does not make this assumption (i.e. model misspecification is acknowledged), and updates the posterior distribution $\pi_n(F)$ in a nonparametric fashion. This in turn leads to more robust posterior inferences on $\theta^*$, as well as asymptotically superior predictions. This improvement in prediction is indeed observed practically as well [Fong et al. (2019)].
We now outline what we mean specifically by superior asymptotic posteriors and predictions. Lyddon et al. (2018) show that the Bayesian bootstrap posterior (which has the same limit as NPL) asymptotically has the sandwich covariance matrix, which is known to be robust. On the other hand, the parametric posterior does not obtain this variance asymptotically.
For prediction, we are interested in the posterior predictive density,
$p_n(y) = \int f_\theta(y) \pi_n(\theta)$, where $\pi_n$ is either the Bayesian or the NPL posterior. Theorem 1 of Lyddon et al. (2018) shows that asymptotically, the KL divergence between $F^*$ and $P_n$ is smaller for the NPL posterior compared to the Bayesian posterior. This asymptotic improvement is indeed due to the robust sandwich covariance matrix [Muller (2013)].
Thank you again for highlighting this important point. We will add a detailed discussion of how NPL addresses model misspecification in the revision paper. And if there are no additional questions or uncertainties, we kindly request you to contemplate revisiting your evaluation to ensure it accurately reflects the situation.
__References__
[1] Walker, S. G. (2013). Bayesian inference with misspecified models. Journal of statistical planning and inference, 143(10), 1621-1633.
[2] Lyddon, S. P., Holmes, C. C., & Walker, S. G. (2019). General Bayesian updating and the loss-likelihood bootstrap. Biometrika, 106(2), 465-478.
[3] Lyddon, S., Walker, S., & Holmes, C. C. (2018). Nonparametric learning from Bayesian models with randomized objective functions. Advances in neural information processing systems, 31.
[4] Müller, U. K. (2013). Risk of Bayesian inference in misspecified models, and the sandwich covariance matrix. Econometrica, 81(5), 1805-1849.
[5] Fong, E., Lyddon, S., & Holmes, C. (2019, May). Scalable nonparametric sampling from multimodal posteriors with the posterior bootstrap. In International Conference on Machine Learning (pp. 1952-1962). PMLR.