Stochastic Latent Residual Video Prediction

Designing video prediction models that account for the inherent uncertainty\nof the future is challenging. Most works in the literature are based on\nstochastic image-autoregressive recurrent networks, which raises several\nperformance and applicability issues. An alternative is to use fully latent\ntemporal models which untie frame synthesis and temporal dynamics. However, no\nsuch model for stochastic video prediction has been proposed in the literature\nyet, due to design and training difficulties. In this paper, we overcome these\ndifficulties by introducing a novel stochastic temporal model whose dynamics\nare governed in a latent space by a residual update rule. This first-order\nscheme is motivated by discretization schemes of differential equations. It\nnaturally models video dynamics as it allows our simpler, more interpretable,\nlatent model to outperform prior state-of-the-art methods on challenging\ndatasets.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC