Model Selection for Bayesian Autoencoders

We develop a novel method for carrying out model selection for Bayesian\nautoencoders (BAEs) by means of prior hyper-parameter optimization. Inspired by\nthe common practice of type-II maximum likelihood optimization and its\nequivalence to Kullback-Leibler divergence minimization, we propose to optimize\nthe distributional sliced-Wasserstein distance (DSWD) between the output of the\nautoencoder and the empirical data distribution. The advantages of this\nformulation are that we can estimate the DSWD based on samples and handle\nhigh-dimensional problems. We carry out posterior estimation of the BAE\nparameters via stochastic gradient Hamiltonian Monte Carlo and turn our BAE\ninto a generative model by fitting a flexible Dirichlet mixture model in the\nlatent space. Consequently, we obtain a powerful alternative to variational\nautoencoders, which are the preferred choice in modern applications of\nautoencoders for representation learning with uncertainty. We evaluate our\napproach qualitatively and quantitatively using a vast experimental campaign on\na number of unsupervised learning tasks and show that, in small-data regimes\nwhere priors matter, our approach provides state-of-the-art results,\noutperforming multiple competitive baselines.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC