Response to Review SQRJ (Part 2/2)
(continued from Part 1/2)
**Q3: User-centered metrics such as perceived realism.**
**A3:** Appreciate this comment. As suggested, in addition to widely adopted metric FID for image realism evaluation, we include a new user-centered metric to access image realism in user study. In this way, we have three criteria in user study: 1) preservation of the fine-grained appearance details, 2) preservation of the garment contour across diverse human body shapes, 3) the perceived image realism. We show the additional comparison results of image realism evaluation in the attached [Fig. D](https://github.com/abcd1905/iclr2025-spm/blob/a3a2f0a84d1ec31d4795b0933946a3f3102f850f/figD.pdf), which again demonstrates the superiority of our SPM-Diff against all baselines in terms of image realism. These results and discussion will be added in revised version.
**Q4: Contribution; warp-based methods.**
**A4:** We appreciate the suggested reference WarpDiffusion [5], and we are happy to discuss the differences between our SPM-Diff and these existing warp-based methods [4, 5]. In particular, existing warp-based methods [4, 5] commonly adopt warping model to warp the input garment according to the input person image, and then directly leverage the warped garment image as **hard/strong condition** to generate VTON results. This way heavily relies on the accuracy of garment warping, thereby easily resulting in unsatisfactory VTON results given the inaccurate warped garments under challenging human poses (see the results of GP-VTON in [Fig. E](https://github.com/abcd1905/iclr2025-spm/blob/a3a2f0a84d1ec31d4795b0933946a3f3102f850f/figE.pdf)). On the contrary, our proposed SPM-Diff adopts a soft way to exploit the visual correspondence prior learnt via the warping model as a **soft condition** to boost VTON. Technically, our SPM-Diff injects the local semantic point features of the input garment into the corresponding positions on the human body according to a flow map estimated by the warping model. This way nicely sidesteps the inputs of holistic warped garment image with amplified pixel-level projection errors, and only emphasizes the visual correspondence of the most informative semantic points between in-shop garment and output synthetic person image. Note that such visual correspondence of local semantic points are exploited in latent space (corresponding to each local region), instead of precise pixel-level location. Thus when encountering mismatched points with slight displacements within local region, SPM-Diff still leads to similar geometry/garment features as matched points, making the visual correspondence more noise-resistant (i.e., more robust to warping results). As shown in [Fig. E](https://github.com/abcd1905/iclr2025-spm/blob/a3a2f0a84d1ec31d4795b0933946a3f3102f850f/figE.pdf), given warping results with severe distoration, our SPM-Diff still manages to achieve promising results, which basically validate the effectiveness of our proposal. We will add the results and discussion in revision.
References:
[1] Zhang, Hongwen and Tian, Yating and Zhang, Yuxiang and Li, Mengcheng and An, Liang and Sun, Zhenan and Liu, Yebin. Pymaf-x: Towards well-aligned full-body model regression from monocular images. In IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
[2] Lin, Jing and Zeng, Ailing and Wang, Haoqian and Zhang, Lei and Li, Yu. One-stage 3d whole-body mesh recovery with component aware transformer. In CVPR, 2023.
[3] Goel, Shubham and Pavlakos, Georgios and Rajasegaran, Jathushan and Kanazawa, Angjoo and Malik, Jitendra. Humans in 4D: Reconstructing and tracking humans with transformers. In ICCV, 2023.
[4] Xie, Zhenyu and Huang, Zaiyu and Dong, Xin and Zhao, Fuwei and Dong, Haoye and Zhang, Xijin and Zhu, Feida and Liang, Xiaodan. Gp-vton: Towards general purpose virtual try-on via collaborative local-flow global-parsing learning. In CVPR, 2023.
[5] Li, Xiu and Kampffmeyer, Michael and Dong, Xin and Xie, Zhenyu and Zhu, Feida and Dong, Haoye and Liang, Xiaodan and others. WarpDiffusion: Efficient Diffusion Model for High-Fidelity Virtual Try-on. In arXiv preprint arXiv:2312.03667.