An Image is Worth More Than a Thousand Words: Towards Disentanglement in the Wild

Unsupervised disentanglement has been shown to be theoretically impossible\nwithout inductive biases on the models and the data. As an alternative\napproach, recent methods rely on limited supervision to disentangle the factors\nof variation and allow their identifiability. While annotating the true\ngenerative factors is only required for a limited number of observations, we\nargue that it is infeasible to enumerate all the factors of variation that\ndescribe a real-world image distribution. To this end, we propose a method for\ndisentangling a set of factors which are only partially labeled, as well as\nseparating the complementary set of residual factors that are never explicitly\nspecified. Our success in this challenging setting, demonstrated on synthetic\nbenchmarks, gives rise to leveraging off-the-shelf image descriptors to\npartially annotate a subset of attributes in real image domains (e.g. of human\nfaces) with minimal manual effort. Specifically, we use a recent language-image\nembedding model (CLIP) to annotate a set of attributes of interest in a\nzero-shot manner and demonstrate state-of-the-art disentangled image\nmanipulation results.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC