In this paper, we investigate the degree to which the encoding of a\n$\\beta$-VAE captures label information across multiple architectures on Binary\nStatic MNIST and Omniglot. Even though they are trained in a completely\nunsupervised manner, we demonstrate that a $\\beta$-VAE can retain a large\namount of label information, even when asked to learn a highly compressed\nrepresentation.\n