Unsupervised Distillation of Syntactic Information from Contextualized Word Representations

Contextualized word representations, such as ELMo and BERT, were shown to\nperform well on various semantic and syntactic tasks. In this work, we tackle\nthe task of unsupervised disentanglement between semantics and structure in\nneural language representations: we aim to learn a transformation of the\ncontextualized vectors, that discards the lexical semantics, but keeps the\nstructural information. To this end, we automatically generate groups of\nsentences which are structurally similar but semantically different, and use\nmetric-learning approach to learn a transformation that emphasizes the\nstructural component that is encoded in the vectors. We demonstrate that our\ntransformation clusters vectors in space by structural properties, rather than\nby lexical semantics. Finally, we demonstrate the utility of our distilled\nrepresentations by showing that they outperform the original contextualized\nrepresentations in a few-shot parsing setting.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC