Unsupervised Disentanglement of Pose, Appearance and Background from Images and Videos

Unsupervised landmark learning is the task of learning semantic keypoint-like\nrepresentations without the use of expensive input keypoint-level annotations.\nA popular approach is to factorize an image into a pose and appearance data\nstream, then to reconstruct the image from the factorized components. The pose\nrepresentation should capture a set of consistent and tightly localized\nlandmarks in order to facilitate reconstruction of the input image. Ultimately,\nwe wish for our learned landmarks to focus on the foreground object of\ninterest. However, the reconstruction task of the entire image forces the model\nto allocate landmarks to model the background. This work explores the effects\nof factorizing the reconstruction task into separate foreground and background\nreconstructions, conditioning only the foreground reconstruction on the\nunsupervised landmarks. Our experiments demonstrate that the proposed\nfactorization results in landmarks that are focused on the foreground object of\ninterest. Furthermore, the rendered background quality is also improved, as the\nbackground rendering pipeline no longer requires the ill-suited landmarks to\nmodel its pose and appearance. We demonstrate this improvement in the context\nof the video-prediction task.\n

Paper

References (44)

Scroll for more · 32 remaining

Similar papers

© 2026 NYSGPT2525 LLC