A Contrastive Learning Approach for Training Variational Autoencoder Priors

Variational autoencoders (VAEs) are one of the powerful likelihood-based\ngenerative models with applications in many domains. However, they struggle to\ngenerate high-quality images, especially when samples are obtained from the\nprior without any tempering. One explanation for VAEs' poor generative quality\nis the prior hole problem: the prior distribution fails to match the aggregate\napproximate posterior. Due to this mismatch, there exist areas in the latent\nspace with high density under the prior that do not correspond to any encoded\nimage. Samples from those areas are decoded to corrupted images. To tackle this\nissue, we propose an energy-based prior defined by the product of a base prior\ndistribution and a reweighting factor, designed to bring the base closer to the\naggregate posterior. We train the reweighting factor by noise contrastive\nestimation, and we generalize it to hierarchical VAEs with many latent variable\ngroups. Our experiments confirm that the proposed noise contrastive priors\nimprove the generative performance of state-of-the-art VAEs by a large margin\non the MNIST, CIFAR-10, CelebA 64, and CelebA HQ 256 datasets. Our method is\nsimple and can be applied to a wide variety of VAEs to improve the expressivity\nof their prior distribution.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC