Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation

Image diffusion distillation achieves high-fidelity generation with very few sampling steps. However, applying these techniques directly to video diffusion often results in unsatisfactory frame quality due to the limited visual quality in public video datasets. This affects the performance of both teacher and student video diffusion models. Our study aims to improve video diffusion distillation while improving frame appearance using abundant high-quality image data. We propose motion consistency model (MCM), a single-stage video diffusion distillation method that disentangles motion and appearance learning. Specifically, MCM includes a video consistency model that distills motion from the video teacher model, and an image discriminator that enhances frame appearance to match high-quality image data. This combination presents two challenges: (1) conflicting frame learning objectives, as video distillation learns from low-quality video frames while the image discriminator targets high-quality images; and (2) training-inference discrepancies due to the differing quality of video samples used during training and inference. To address these challenges, we introduce disentangled motion distillation and mixed trajectory distillation. The former applies the distillation objective solely to the motion representation, while the latter mitigates training-inference discrepancies by mixing distillation trajectories from both the low- and high-quality video domains. Extensive experiments show that our MCM achieves the state-of-the-art video diffusion distillation performance. Additionally, our method can enhance frame quality in video diffusion models, producing frames with high aesthetic scores or specific styles without corresponding video data.

Paper

References (72)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer sR1C6/10 · confidence 3/52024-07-08

Summary

This work presented a new consistency based framework for video diffusion model distillation. Specifically, the adversarial loss is leveraged to enhance the video quality and consistency distillation loss is performed in the motion embedding space to learn the video motion patterns effectively. In addition, the authors proposed mixed trajectory distillation to ensure better alignment between training and inference phases. The experimental results demonstrate the proposed approach could produce more visual-pleasing results compared with previous distillation methods.

Strengths

1. The proposed disentangled motion-appearance distillation is reasonable and effective. 2. The generated results in Fig. 5 and supp are very promising. 3. The quantitive comparisons in Table 1 and 2 are convincing.

Weaknesses

1. The adversarial loss is not stable and could the authors employ other manners such as perceptual loss?

Questions

Could the proposed algorithm achieves satisfactory performance on other video generation tasks such as StableVideoDiffusion (Image-to-video) and AnimateAnyone (human video generation)?

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

4

Limitations

The proposed algorithm is not evaluated on high-resolution video generation diffusion models such as (1024x576 or 768x768).

Reviewer FQyx7/10 · confidence 3/52024-07-13

Summary

This paper proposed a single-stage video diffusion distillation method that can disentangle motion and appearance learning, thus improving frame appearance using various high-quality image data. The proposed mixed trajectory distillation mitigates the training-inference differences in terms of video quality. The extensive experiments demonstrate the superior performance in enhancing frame quality in the video diffusion model.

Strengths

1. The proposed disentangled motion distillation and mixed trajectory distillation are intuitive and novel. 2. The experiments are thorough. They are conducted across various datasets and show superior results in terms of video diffusion distillation. The ablation study shows the effectiveness of the proposed disentangled motion distribution and mixed trajectory distribution modules. 3. The paper is well-written and easy to follow.

Weaknesses

1. Motion jittering in the supplementary video. It's probably caused by the teacher model but the authors could better discuss the way to alleviate it. 2. In Fig. 6, there is no caption to indicate which one is the result of the proposed methods and which one is the designed two-stage baseline. What are the differences between the first row and the second row?

Questions

1. In Fig. 6, why do the "Ours w/ Webvid" results also have watermarks?

Rating

7

Confidence

3

Soundness

4

Presentation

3

Contribution

3

Limitations

The authors have discussed the limitations of this paper

Reviewer VLKW5/10 · confidence 5/52024-07-15

Summary

This paper proposes a video diffusion distillation method that disentangles motion and appearance learning. Basically, it proposes to enhance the appearance generation with high-quality image data and distill motion knowledge from the video teacher model.

Strengths

1. The proposed method can distill motion knowledge from video diffusion models and improve the appearance quality through disentangled motion distillation. 2. The mixed trajectory distillation is proposed to improve training-inference alignment and enhance generation quality. 3. This paper is technically clear and the organization is good.

Weaknesses

1. The introduction of gaps between the training and inference distillation inputs is not so straightforward. 2. The introduction to related work needs to be significantly enhanced, especially in terms of the idea of decoupling appearance and motion, which is no longer uncommon and has many related works. 3. It simply provides the conclusion that "learnable representation works the best" without giving specific analysis as to why. Such an analysis may be more helpful for following research.

Questions

Please refer to the weaknesses.

Rating

5

Confidence

5

Soundness

3

Presentation

3

Contribution

2

Limitations

Broader impacts and limitations have been discussed in the paper.

Reviewer VLKW2024-08-12

The authors have addressed most of my concerns. However, the response to Weakness 2 is still very sketchy. I suggest the authors provide more detailed discussions in the revised paper. I decide to keep my initial positive rating.

Reviewer FQyx2024-08-08

I appreciate the author's response. It has addressed my concerns.

Reviewer sR1C2024-08-13

Thanks for the response. My concerns have been addressed well.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC