A Sober Look at the Unsupervised Learning of Disentangled Representations and their Evaluation

The idea behind the \\emph{unsupervised} learning of \\emph{disentangled}\nrepresentations is that real-world data is generated by a few explanatory\nfactors of variation which can be recovered by unsupervised learning\nalgorithms. In this paper, we provide a sober look at recent progress in the\nfield and challenge some common assumptions. We first theoretically show that\nthe unsupervised learning of disentangled representations is fundamentally\nimpossible without inductive biases on both the models and the data. Then, we\ntrain over $14000$ models covering most prominent methods and evaluation\nmetrics in a reproducible large-scale experimental study on eight data sets. We\nobserve that while the different methods successfully enforce properties\n"encouraged" by the corresponding losses, well-disentangled models seemingly\ncannot be identified without supervision. Furthermore, different evaluation\nmetrics do not always agree on what should be considered "disentangled" and\nexhibit systematic differences in the estimation. Finally, increased\ndisentanglement does not seem to necessarily lead to a decreased sample\ncomplexity of learning for downstream tasks. Our results suggest that future\nwork on disentanglement learning should be explicit about the role of inductive\nbiases and (implicit) supervision, investigate concrete benefits of enforcing\ndisentanglement of the learned representations, and consider a reproducible\nexperimental setup covering several data sets.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC