Mutual Information Based Method for Unsupervised Disentanglement of Video Representation

Video Prediction is an interesting and challenging task of predicting future\nframes from a given set context frames that belong to a video sequence. Video\nprediction models have found prospective applications in Maneuver Planning,\nHealth care, Autonomous Navigation and Simulation. One of the major challenges\nin future frame generation is due to the high dimensional nature of visual\ndata. In this work, we propose Mutual Information Predictive Auto-Encoder\n(MIPAE) framework, that reduces the task of predicting high dimensional video\nframes by factorising video representations into content and low dimensional\npose latent variables that are easy to predict. A standard LSTM network is used\nto predict these low dimensional pose representations. Content and the\npredicted pose representations are decoded to generate future frames. Our\napproach leverages the temporal structure of the latent generative factors of a\nvideo and a novel mutual information loss to learn disentangled video\nrepresentations. We also propose a metric based on mutual information gap (MIG)\nto quantitatively access the effectiveness of disentanglement on DSprites and\nMPI3D-real datasets. MIG scores corroborate with the visual superiority of\nframes predicted by MIPAE. We also compare our method quantitatively on\nevaluation metrics LPIPS, SSIM and PSNR.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC