Response to Review gZ41
1 Results on Video Frame Interpolation with original key frames.
Thanks for pointing out this important evaluation setting. Following your valuable suggestions, we conduct a quantitative and qualitative evaluation of the performance of MaskINT in the reconstructive setting, where it engages in video frame interpolation using the original key frames. In this evaluation, we apply signal-to-noise ratio (PSNR), learned perceptual image patch similarity (LPIPS), and structured similarity (SSIM) to compare the interpolated frames with original video frames. We benchmark our method against two state-of-the-art Video Frame Interpolation (VFI) methods, including FILM [A] and RIFE [B]. Table 3 and Figure 6 in Appendix show that our method significantly outperforms VFI methods on all evaluation metrics, with the benefit of the structure guidance from the intermediate frames (For better visualization, we suggest you watch the .mp4 video in Supplementary).
Moreover, even when confronted with significant motion between two frames, our approach successfully reconstructs the original video, maintaining consistent motion through the aid of structural guidance. In contrast, FILM introduces undesirable artifacts, including distorted background, multiple cat hands, and the absence of a camel's head, etc. The major reason is that current VFI models mainly focus on generating slow-motion effects and enhancing frame rate, making them less effective in handling frames with large motions. Additionally, the absence of structural guidance poses a challenge for these methods in accurately aligning generated videos with the original motion.
2 Comparisons with [1] and [2] with additional structure control.
Thanks for sharing these important works. However, We would like to respectfully point out that designing an architecture to explicitly introduce structural control to these works is a non-trivial task.Besides, these methods are not open-source and they use some in-house datasets to train the model, making it difficult to make any further adjustments and conduct a fair comparison.
3 Performance with single frame.
We further evaluate the performance where only the initial edited key frame is given. As shown in Table 2, when there is only one keyframe, the performance is downgraded since it’s difficult to understand the motion within the video with a single reference frame.
[A] Reda, Fitsum, et al. "Film: Frame interpolation for large motion." European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022.
[B] Huang, Zhewei, et al. "Real-time intermediate flow estimation for video frame interpolation." European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022