Textual Semantics Matters: Unsupervised Representation Disentanglement in Realistic Scenarios with Language Inductive Bias
Representation disentanglement helps AI-fundamentally understand the real world in an interpretable manner. Existing studies of disentangled representation learning (DRL) are typically explored in the image domain only with visual constraints. They largely neglect the essence that discrete-valued semantics denoted by text are originally disentangled, which contains critical properties promoting the disentanglement. In this work, we comprehensively investigate the discrete semantics in the text domain and propose to enhance the DRL model with three types of underlying language inductive bias-image-text alignment, cycle consistency, and shift equivariant. Specifically, we propose to set up the DRL workflow using a diffusion procedure, which allows us to apply autoregressive models in a continuous-valued image space, with discrete-valued text constraints. Furthermore, instead of relying solely on the reconstruction loss, we define a set of Text-Aided Regularization to encourage representation disentanglement. In addition, an auxiliary condition branch is introduced to employ text regularization, enabling constraints over global data with efficient fine-tuning. Experiments demonstrate that textual semantics matters and the proposed text-aided regularization is indispensable towards a stronger DRL framework, especially in non-toy/-synthetic realistic scenarios.
Paper
Full text
Textual Semantics Matters: Unsupervised Representation Disentanglement in Realistic Scenarios with Language Inductive Bias
Semantic Scholar · Computer Science · 2025
Abstract
Representation disentanglement helps AI-fundamentally understand the real world in an interpretable manner. Existing studies of disentangled representation learning (DRL) are typically explored in the image domain only with visual constraints. They largely neglect the essence that discrete-valued semantics denoted by text are originally disentangled, which contains critical properties promoting the disentanglement. In this work, we comprehensively investigate the discrete semantics in the text domain and propose to enhance the DRL model with three types of underlying language inductive bias-image-text alignment, cycle consistency, and shift equivariant. Specifically, we propose to set up the DRL workflow using a diffusion procedure, which allows us to apply autoregressive models in a continuous-valued image space, with discrete-valued text constraints. Furthermore, instead of relying solely on the reconstruction loss, we define a set of Text-Aided Regularization to encourage representation disentanglement. In addition, an auxiliary condition branch is introduced to employ text regularization, enabling constraints over global data with efficient fine-tuning. Experiments demonstrate that textual semantics matters and the proposed text-aided regularization is indispensable towards a stronger DRL framework, especially in non-toy/-synthetic realistic scenarios.