In many data analysis tasks, it is beneficial to learn representations where\neach dimension is statistically independent and thus disentangled from the\nothers. If data generating factors are also statistically independent,\ndisentangled representations can be formed by Bayesian inference of latent\nvariables. We examine a generalization of the Variational Autoencoder (VAE),\n$\\beta$-VAE, for learning such representations using variational inference.\n$\\beta$-VAE enforces conditional independence of its bottleneck neurons\ncontrolled by its hyperparameter $\\beta$. This condition is in general not\ncompatible with the statistical independence of latents. By providing\nanalytical and numerical arguments, we show that this incompatibility leads to\na non-monotonic inference performance in $\\beta$-VAE with a finite optimal\n$\\beta$.\n