Latent Image Animator: Learning to Animate Images via Latent Space Navigation

Due to the remarkable progress of deep generative models, animating images\nhas become increasingly efficient, whereas associated results have become\nincreasingly realistic. Current animation-approaches commonly exploit structure\nrepresentation extracted from driving videos. Such structure representation is\ninstrumental in transferring motion from driving videos to still images.\nHowever, such approaches fail in case the source image and driving video\nencompass large appearance variation. Moreover, the extraction of structure\ninformation requires additional modules that endow the animation-model with\nincreased complexity. Deviating from such models, we here introduce the Latent\nImage Animator (LIA), a self-supervised autoencoder that evades need for\nstructure representation. LIA is streamlined to animate images by linear\nnavigation in the latent space. Specifically, motion in generated video is\nconstructed by linear displacement of codes in the latent space. Towards this,\nwe learn a set of orthogonal motion directions simultaneously, and use their\nlinear combination, in order to represent any displacement in the latent space.\nExtensive quantitative and qualitative analysis suggests that our model\nsystematically and significantly outperforms state-of-art methods on VoxCeleb,\nTaichi and TED-talk datasets w.r.t. generated quality.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC