Bayesian neural networks that incorporate data augmentation implicitly use a\n``randomly perturbed log-likelihood [which] does not have a clean\ninterpretation as a valid likelihood function'' (Izmailov et al. 2021). Here,\nwe provide several approaches to developing principled Bayesian neural networks\nincorporating data augmentation. We introduce a ``finite orbit'' setting which\nallows likelihoods to be computed exactly, and give tight multi-sample bounds\nin the more usual ``full orbit'' setting. These models cast light on the origin\nof the cold posterior effect. In particular, we find that the cold posterior\neffect persists even in these principled models incorporating data\naugmentation. This suggests that the cold posterior effect cannot be dismissed\nas an artifact of data augmentation using incorrect likelihoods.\n
Paper
References (66)
Scroll for more · 38 remaining