Do As I Do: Transferring Human Motion and Appearance between Monocular Videos with Spatial and Temporal Constraints

Creating plausible virtual actors from images of real actors remains one of\nthe key challenges in computer vision and computer graphics. Marker-less human\nmotion estimation and shape modeling from images in the wild bring this\nchallenge to the fore. Although the recent advances on view synthesis and\nimage-to-image translation, currently available formulations are limited to\ntransfer solely style and do not take into account the character's motion and\nshape, which are by nature intermingled to produce plausible human forms. In\nthis paper, we propose a unifying formulation for transferring appearance and\nretargeting human motion from monocular videos that regards all these aspects.\nOur method synthesizes new videos of people in a different context where they\nwere initially recorded. Differently from recent appearance transferring\nmethods, our approach takes into account body shape, appearance, and motion\nconstraints. The evaluation is performed with several experiments using\npublicly available real videos containing hard conditions. Our method is able\nto transfer both human motion and appearance outperforming state-of-the-art\nmethods, while preserving specific features of the motion that must be\nmaintained (e.g., feet touching the floor, hands touching a particular object)\nand holding the best visual quality and appearance metrics such as Structural\nSimilarity (SSIM) and Learned Perceptual Image Patch Similarity (LPIPS).\n

Paper

References (41)

Scroll for more · 29 remaining

Similar papers

© 2026 NYSGPT2525 LLC