Identifiability of a statistical model with two latent vectors: Importance of the dimensionality relation and application to graph embedding
Identifiability of statistical models is a key notion in unsupervised representation learning. Recent work of nonlinear independent component analysis (ICA) employs auxiliary data in addition to input data which is supposed to be observed as a nonlinear mixing of a single latent vector, and has established identifiability conditions. This paper proposes a statistical model of two latent vectors with single auxiliary data generalizing nonlinear ICA, and establishes various identifiability conditions. Unlike previous work, the two latent vectors in the proposed model can have different dimensions, and this property enables us to reveal an insightful dimensionality condition among two latent vectors and auxiliary data. Furthermore, surprisingly, we prove that the proposed model with the nonlinear mixing has the same indeterminacies as linear ICA under certain conditions: The elements in the latent vector can be recovered up to their permutation and scales. Next, we apply the identifiability theory to a statistical model for graph data. As a result, one of the identifiability conditions includes an appealing implication: Identifiability of the statistical model could depend on the maximum value of link weights in graph data. Then, we propose a practical method for identifiable graph embedding. Finally, we numerically demonstrate that the proposed method well-recovers the latent vectors and identifiability clearly depends on the maximum value of link weights, which supports the implication of our theoretical results.
Paper
References (48)
Scroll for more · 36 remaining