Energy-based models (EBMs) are generative models that are usually trained via\nmaximum likelihood estimation. This approach becomes challenging in generic\nsituations where the trained energy is non-convex, due to the need to sample\nthe Gibbs distribution associated with this energy. Using general Fenchel\nduality results, we derive variational principles dual to maximum likelihood\nEBMs with shallow overparametrized neural network energies, both in the\nfeature-learning and lazy linearized regimes. In the feature-learning regime,\nthis dual formulation justifies using a two time-scale gradient ascent-descent\n(GDA) training algorithm in which one updates concurrently the particles in the\nsample space and the neurons in the parameter space of the energy. We also\nconsider a variant of this algorithm in which the particles are sometimes\nrestarted at random samples drawn from the data set, and show that performing\nthese restarts at every iteration step corresponds to score matching training.\nThese results are illustrated in simple numerical experiments, which indicates\nthat GDA performs best when features and particles are updated using similar\ntime scales.\n
Paper
References (70)
Scroll for more · 38 remaining