Unsupervised learning of disentangled representations in deep restricted kernel machines with orthogonality constraints
We introduce Constr-DRKM, a deep kernel method for the unsupervised learning\nof disentangled data representations. We propose augmenting the original deep\nrestricted kernel machine formulation for kernel PCA by orthogonality\nconstraints on the latent variables to promote disentanglement and to make it\npossible to carry out optimization without first defining a stabilized\nobjective. After illustrating an end-to-end training procedure based on a\nquadratic penalty optimization algorithm with warm start, we quantitatively\nevaluate the proposed method's effectiveness in disentangled feature learning.\nWe demonstrate on four benchmark datasets that this approach performs similarly\noverall to $\\beta$-VAE on a number of disentanglement metrics when few training\npoints are available, while being less sensitive to randomness and\nhyperparameter selection than $\\beta$-VAE. We also present a deterministic\ninitialization of Constr-DRKM's training algorithm that significantly improves\nthe reproducibility of the results. Finally, we empirically evaluate and\ndiscuss the role of the number of layers in the proposed methodology, examining\nthe influence of each principal component in every layer and showing that\ncomponents in lower layers act as local feature detectors capturing the broad\ntrends of the data distribution, while components in deeper layers use the\nrepresentation learned by previous layers and more accurately reproduce\nhigher-level features.\n