Learning expressive probabilistic models correctly describing the data is a\nubiquitous problem in machine learning. A popular approach for solving it is\nmapping the observations into a representation space with a simple joint\ndistribution, which can typically be written as a product of its marginals --\nthus drawing a connection with the field of nonlinear independent component\nanalysis. Deep density models have been widely used for this task, but their\nmaximum likelihood based training requires estimating the log-determinant of\nthe Jacobian and is computationally expensive, thus imposing a trade-off\nbetween computation and expressive power. In this work, we propose a new\napproach for exact training of such neural networks. Based on relative\ngradients, we exploit the matrix structure of neural network parameters to\ncompute updates efficiently even in high-dimensional spaces; the computational\ncost of the training is quadratic in the input size, in contrast with the cubic\nscaling of naive approaches. This allows fast training with objective functions\ninvolving the log-determinant of the Jacobian, without imposing constraints on\nits structure, in stark contrast to autoregressive normalizing flows.\n
Paper
References (57)
Scroll for more · 38 remaining