MotionCraft: Physics-based Zero-Shot Video Generation

Generating videos with realistic and physically plausible motion is one of the main recent challenges in computer vision. While diffusion models are achieving compelling results in image generation, video diffusion models are limited by heavy training and huge models, resulting in videos that are still biased to the training dataset. In this work we propose MotionCraft, a new zero-shot video generator to craft physics-based and realistic videos. MotionCraft is able to warp the noise latent space of an image diffusion model, such as Stable Diffusion, by applying an optical flow derived from a physics simulation. We show that warping the noise latent space results in coherent application of the desired motion while allowing the model to generate missing elements consistent with the scene evolution, which would otherwise result in artefacts or missing content if the flow was applied in the pixel space. We compare our method with the state-of-the-art Text2Video-Zero reporting qualitative and quantitative improvements, demonstrating the effectiveness of our approach to generate videos with finely-prescribed complex motion dynamics. Project page: https://mezzelfo.github.io/MotionCraft/

Paper

Similar papers

Peer review

Reviewer zFVb5/10 · confidence 4/52024-06-18

Summary

In this work the authors propose MotionCraft, a new zero-shot video generator to craft physics-based and realistic videos. MotionCraft is able to warp the noise latent space of an image diffusion model, such as Stable Diffusion, by applying an optical flow derived from a physics simulation. The authors show that warping the noise latent space results in coherent application of the desired motion while allowing the model to generate missing elements consistent with the scene evolution, which would otherwise result in artefacts or missing content if the flow was applied in the pixel space. The authors compare the method with the state-of-the-art Text2Video-Zero reporting qualitative and quantitative improvements, demonstrating the effectiveness of the approach to generate videos with finely-prescribed complex motion dynamics.

Strengths

1. This method does not need extra training, which is efficient. 2. This method introduces explicit phisyics control to the field of video generation, which is novel. 3. This method finds out that the optical flow is consistent between pixel space and latent space, which is interesting.

Weaknesses

1. The main concern is about the experiments. This paper only have 5 video results in total, which is not sufficient. I am worrying that the method is highly unstable and not robust, thus the author cannot present more video results. If this is the case, I think this manuscript is not suitable for publication. If this is not the case, I think the authors should provide more generated results and I will be glad to raise my score. 2. How long does it take to generate a single video? If the paper claims that the video can be generated within minutes, I think generating tens of videos in the supplemtary material would be a good idea.

Questions

See the weakness.

Rating

5

Confidence

4

Soundness

2

Presentation

2

Contribution

3

Limitations

See the weakness.

Reviewer zFVb2024-08-11

Thank you for the rebuttal. I have read the additional results. The rating has been updated. I hope the authors could show more generated results in the suppmentary materials upon the acception of this manuscript.

Reviewer cTbz5/10 · confidence 4/52024-07-06

Summary

The paper presents MOTIONCRAFT, a novel zero-shot video generation method that leverages physical simulations to create realistic and physically plausible videos. Unlike traditional video diffusion models that require extensive training and large datasets, MOTIONCRAFT uses a pre-trained image diffusion model, such as Stable Diffusion, and warps its noise latent space with optical flow derived from physical simulations. This approach ensures coherent motion and the generation of missing elements consistent with scene evolution. **Key Contributions:** 1. **Innovative Approach:** Introduction of a zero-shot video generation method that uses optical flow from physical simulations to warp the noise latent space of a pre-trained image diffusion model. 2. **Experimental Validation:** Demonstrates the effectiveness of MOTIONCRAFT through both qualitative and quantitative comparisons with the state-of-the-art Text2Video-Zero method, showing significant improvements. 3. **Theoretical Insights:** Provides an analysis of the correlation between optical flow in the image space and the noise latent space, supporting the proposed method. 4. **Versatility:** Showcases the ability of MOTIONCRAFT to generate videos with complex dynamics, including fluid dynamics, rigid body physics, and multi-agent interaction models, without additional training. 5. **Technical Details:** Describes key techniques such as multi-frame cross-attention and spatial noise map weighting to ensure temporal and spatial consistency in the generated videos. Overall, MOTIONCRAFT represents a significant advancement in zero-shot video generation, combining the strengths of physical simulations and image diffusion models to produce high-quality, dynamic videos.

Strengths

Refer to Summary.

Weaknesses

- There are now many approaches to zero-shot video generation, such as [https://openreview.net/forum?id=zOjW6yVYkE](https://openreview.net/forum?id=zOjW6yVYkE). The authors only compared their method with T2V0 (a relatively earlier method), which may make the experimental results insufficient. It would be better to include a more comprehensive comparison. - The current zero-shot video generation methods generally cannot achieve a very coherent video effect and can only generate keyframes. Although the method proposed in the paper largely ensures content consistency between consecutive frames, the generated videos still fail to achieve a highly coherent effect as seen in the demonstration. - Optical flow-based strategies often face limitations in certain specific situations. In complex environments, optical flow might not be effective. This aspect should be discussed and analyzed in the paper's main text. Moreover, the performance in different scenarios should be thoroughly evaluated using a variety of experimental results presented in the paper, instead of merely showcasing the method's validity through three handpicked examples.

Questions

- Is it possible to achieve better controllable video generation by handling optical flow in a manner similar to that mentioned in the Generative Image Dynamics ([https://generative-dynamics.github.io/](https://generative-dynamics.github.io/)) paper? - Can more coherent videos be generated through interpolation in optical flow? - Is it possible to generate long videos?

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

Refer to Weakness

Reviewer ypn76/10 · confidence 4/52024-07-14

Summary

The paper works on the zero-shot video generation task and proposes, MotionCraft. It uses physics simulations to generate optical flow that follows physical dynamics. Then, optical flow is applied to warp the noise in the latent space with the stable diffusion model. This approach ensures coherent motion application and consistent scene evolution, avoiding artefacts and missing content typical in pixel space flow applications. Compared to the state-of-the-art Text2Video-Zero, MotionCraft shows both qualitative and quantitative improvements in generating videos with complex motion dynamics.

Strengths

1. The paper is well-motivated and well-written. 2. The idea of using physics simulator to generate the optical flow which is then applied in latent space is very interesting. 3. The qualitative results are impressive.

Weaknesses

1. More quantitative comparison with baselines. Table 1 only reports the comparison with T2V0 on the generated videos. However, it is not clear which benchmark it is. Is it possible to compare with other baselines on more benchmarks, like MUG, MHAD? 2. The method seems limited by specific types of dynamics, like fluid dynamics. It is not clear how to generate more general dynamics in real world. This may limit the potential application of the proposed method. 3. The method is assumes that "Optical Flow is preserved in the Latent Space of Stable Diffusion" based on the observation of average correlations 0.727 between optical flows estimated in the RGB and noise latent spaces. Does this assumption hold true for generating realistic, pixel-wise precise motion in video with only 0.72 cosine simlarity?

Questions

How do you specify the region for simulating physical dynamics? Will the type of dynamic physics simulator affect the quality of the generated video?

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The method is limited to specific types of dynamic simulators and may be hard to apply to generic real-world video generation.

Area Chair FjZe2024-08-12

Please discuss

Dear reviewer, The discussion period is coming to a close soon. Please do your best to engage with the authors. Thank you, Your AC

Area Chair FjZe2024-08-12

Please discuss

Dear reviewer, The discussion period is coming to a close soon. Please do your best to engage with the authors. Thank you, Your AC

Reviewer cTbz2024-08-14

Thank you for your rebuttal. The limited results remain a major concern. I will maintain my original rating.

Authorsrebuttal2024-08-14

We respectfully disagree with the reviewer regarding the perceived limitation in the number of examples shown. In the original manuscript, we included 6 examples (5 in the main text and 1 in the appendix). With the rebuttal, we added 9 more examples, effectively more than doubling the number of experiments. Overall, the examples provided highlight the capacity of MotionCraft to handle a variety of scenarios, including different physics (rigid bodies, fluids, multi-agents), different diffusion models with varying resolutions (Stable Diffusion and SDXL), and different frame rates (even reaching videos of 200 frames). Additionally, we would like to point out that the number of compatible examples provided by T2V0 (those without trained ControlNets) is comparable to the number of examples we have presented.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC