Vid-ODE: Continuous-Time Video Generation with Neural Ordinary Differential Equation

Video generation models often operate under the assumption of fixed frame\nrates, which leads to suboptimal performance when it comes to handling flexible\nframe rates (e.g., increasing the frame rate of the more dynamic portion of the\nvideo as well as handling missing video frames). To resolve the restricted\nnature of existing video generation models' ability to handle arbitrary\ntimesteps, we propose continuous-time video generation by combining neural ODE\n(Vid-ODE) with pixel-level video processing techniques. Using ODE-ConvGRU as an\nencoder, a convolutional version of the recently proposed neural ODE, which\nenables us to learn continuous-time dynamics, Vid-ODE can learn the\nspatio-temporal dynamics of input videos of flexible frame rates. The decoder\nintegrates the learned dynamics function to synthesize video frames at any\ngiven timesteps, where the pixel-level composition technique is used to\nmaintain the sharpness of individual frames. With extensive experiments on four\nreal-world video datasets, we verify that the proposed Vid-ODE outperforms\nstate-of-the-art approaches under various video generation settings, both\nwithin the trained time range (interpolation) and beyond the range\n(extrapolation). To the best of our knowledge, Vid-ODE is the first work\nsuccessfully performing continuous-time video generation using real-world\nvideos.\n

Paper

References (34)

Scroll for more · 22 remaining

Similar papers

© 2026 NYSGPT2525 LLC