Learning to predict the long-term future of video frames is notoriously\nchallenging due to inherent ambiguities in the distant future and dramatic\namplifications of prediction error through time. Despite the recent advances in\nthe literature, existing approaches are limited to moderately short-term\nprediction (less than a few seconds), while extrapolating it to a longer future\nquickly leads to destruction in structure and content. In this work, we revisit\nhierarchical models in video prediction. Our method predicts future frames by\nfirst estimating a sequence of semantic structures and subsequently translating\nthe structures to pixels by video-to-video translation. Despite the simplicity,\nwe show that modeling structures and their dynamics in the discrete semantic\nstructure space with a stochastic recurrent estimator leads to surprisingly\nsuccessful long-term prediction. We evaluate our method on three challenging\ndatasets involving car driving and human dancing, and demonstrate that it can\ngenerate complicated scene structures and motions over a very long time horizon\n(i.e., thousands frames), setting a new standard of video prediction with\norders of magnitude longer prediction time than existing approaches. Full\nvideos and codes are available at https://1konny.github.io/HVP/.\n
Paper
References (52)
Scroll for more · 38 remaining