Ada-VE: Training-Free Consistent Video Editing Using Adaptive Motion Prior

Video-to- video synthesis poses significant challenges in maintaining character consistency, smooth temporal tran-sitions, and preserving visual quality during fast motion. While recent fully cross-frame self-attention mechanisms have improved character consistency across multiple frames, they come with high computational costs and often include re-dundant operations, especially for videos with higher frame rates. To address these inefficiencies, we propose an adaptive motion-guided cross-frame attention mechanism that selectively reduces redundant computations. This enables a greater number of cross-frame attentions over more frames within the same computational budget, thereby enhancing both video quality and temporal coherence. Our method leverages optical flow to focus on moving regions while sparsely attending to stationary areas, allowing for the joint editing of more frames without increasing computational demands. Traditional frame interpolation techniques struggle with motion blur and flickering in intermediate frames, which compromises visual fidelity. To mitigate this, we intro-duce KV-caching for jointly edited frames, reusing keys and values across intermediate frames to preserve visual quality and maintain temporal consistency throughout the video. With our adaptive cross-frame self-attention approach, we achieve a threefold increase in the number of keyframes processed compared to existing methods, all within the same computational budget as fully cross-frame attention base-lines. This results in significant improvements in prediction accuracy and temporal consistency, outperforming state-of-the-art approaches. Code is made publicly available at https://github.com/tanvir-utexaslAdaVEltree/main.

Paper

References (47)

Scroll for more · 35 remaining

Similar papers

© 2026 NYSGPT2525 LLC