Law of Large Numbers for Bayesian two-layer Neural Network trained with Variational Inference
We provide a rigorous analysis of training by variational inference (VI) of\nBayesian neural networks in the two-layer and infinite-width case. We consider\na regression problem with a regularized evidence lower bound (ELBO) which is\ndecomposed into the expected log-likelihood of the data and the\nKullback-Leibler (KL) divergence between the a priori distribution and the\nvariational posterior. With an appropriate weighting of the KL, we prove a law\nof large numbers for three different training schemes: (i) the idealized case\nwith exact estimation of a multiple Gaussian integral from the\nreparametrization trick, (ii) a minibatch scheme using Monte Carlo sampling,\ncommonly known as Bayes by Backprop, and (iii) a new and computationally\ncheaper algorithm which we introduce as Minimal VI. An important result is that\nall methods converge to the same mean-field limit. Finally, we illustrate our\nresults numerically and discuss the need for the derivation of a central limit\ntheorem.\n
Paper
References (48)
Scroll for more · 36 remaining