Weaknesses
1. The pictures in Figure 1 are too small, especially the optical flow map is not clear enough, no obvious changes can be seen before and after refinement. Also, is the schematic diagram of affine transformation in Figure 1 drawn based on a real example? Why does the affine transformation change so much after refinement?
2. Straightforward combination of existing techniques. The innovation of this paper is not enough, the main innovation point lies in the motion refinement module. However, the method of correlation volume and iteratively refine optical flow used in it is very similar to the correlation matrix and iteratively update in some flow estimation[1,2,3] and correspondence estimation[4,5] methods, and this set of process of first constructing a 4D correlation volume, and then iteratively updating and optimizing optical flow has also been used in some neural style transfer(NST) methods[6,7], but the author did not clarify the differences between the module they used in this paper and related modules in these NST methods.
3. There is an error in line 180 of the article: the author claims that the look up operation on the correlation volume is shown on the left side of Figure 2, but there is no corresponding content in Figure 2. Should Figure 2 be changed to Figure 3?
4. In the comparison video provided in the supplemental material, the video of each method is too small, and it is difficult to see the tiny facial deformation details. It is recommended to arrange the videos of each method in the form of a nine-square grid, and enlarge the size of each video window.
5. The experiments of verifying the effectiveness of the proposed method are insufficient. The framework used by your method is the framework of unsupervised image animation [8,9,10,11], which should be able to generate animation videos on any object category. In addition to human faces, this framework can also be applied to human bodies and animals. Previous unsupervised image animation methods [8,9,10,11] have also been tested on related datasets of different subjects, including TaiChiHD[9], TED-talks[10], and MGif[8]. Your method has only been tested on face-related datasets, but from your overall framework, I don't see any modules that restrict the method to only work on human face. Therefore, I think that the non-prior-based motion refinement module should be applied to datasets of different objects to further test its effectiveness.
6. Lack of quantitative experiments on the Cross-identity task, such as user study that have been used in previous methods [8,9,10,11].
References
[1] Dosovitskiy, Alexey, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick van der Smagt, Daniel Cremers, and Thomas Brox. 2015. “FlowNet: Learning Optical Flow with Convolutional Networks.” In 2015 IEEE International Conference on Computer Vision (ICCV). doi:10.1109/iccv.2015.316.
[2] Teed, Zachary, and Jia Deng. 2020. “RAFT: Recurrent All-Pairs Field Transforms for Optical Flow.” In Computer Vision – ECCV 2020, Lecture Notes in Computer Science, 402–19. doi:10.1007/978-3-030-58536-5_24.
[3] Xu, Haofei, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, and Dacheng Tao. 2022. “GMFlow: Learning Optical Flow via Global Matching.” In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). doi:10.1109/cvpr52688.2022.00795.
[4] Kim, Seungryong, Stephen Lin, SangRyul Jeon, Dongbo Min, and Kwanghoon Sohn. 2018. “Recurrent Transformer Networks for Semantic Correspondence.” arXiv: Computer Vision and Pattern Recognition,arXiv: Computer Vision and Pattern Recognition, October.
[5] Zhang, Pan, Bo Zhang, Dong Chen, Lu Yuan, and Fang Wen. 2020. “Cross-Domain Correspondence Learning for Exemplar-Based Image Translation.” In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). doi:10.1109/cvpr42600.2020.00519.
[6] Liu, Xiaochang, Xuanyi Li, Ming-Ming Cheng, and Peter Hall. 2020. “Geometric Style Transfer.” Cornell University - arXiv,Cornell University - arXiv, July.
[7] Liu, Xiao-Chang, Yong-Liang Yang, and Peter Hall. 2021. “Learning to Warp for Style Transfer.” In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). doi:10.1109/cvpr46437.2021.00370.
[8] Siarohin, Aliaksandr, Stephane Lathuiliere, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. “Animating Arbitrary Objects via Deep Motion Transfer.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). doi:10.1109/cvpr.2019.00248.
[9] Siarohin, Aliaksandr, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. “First Order Motion Model for Image Animation.” Neural Information Processing Systems,Neural Information Processing Systems, January.
[10] Siarohin, Aliaksandr, Oliver J. Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov. 2021. “Motion Representations for Articulated Animation.” In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). doi:10.1109/cvpr46437.2021.01344.
[11] Zhao, Jian, and Hui Zhang. n.d. “Thin-Plate Spline Motion Model for Image Animation.”