Disentangling Multiple Features in Video Sequences using Gaussian Processes in Variational Autoencoders
We introduce MGP-VAE (Multi-disentangled-features Gaussian Processes\nVariational AutoEncoder), a variational autoencoder which uses Gaussian\nprocesses (GP) to model the latent space for the unsupervised learning of\ndisentangled representations in video sequences. We improve upon previous work\nby establishing a framework by which multiple features, static or dynamic, can\nbe disentangled. Specifically we use fractional Brownian motions (fBM) and\nBrownian bridges (BB) to enforce an inter-frame correlation structure in each\nindependent channel, and show that varying this structure enables one to\ncapture different factors of variation in the data. We demonstrate the quality\nof our representations with experiments on three publicly available datasets,\nand also quantify the improvement using a video prediction task. Moreover, we\nintroduce a novel geodesic loss function which takes into account the curvature\nof the data manifold to improve learning. Our experiments show that the\ncombination of the improved representations with the novel loss function enable\nMGP-VAE to outperform the baselines in video prediction.\n
Paper
References (28)
Scroll for more · 16 remaining