Syntactic structure of sentences in a document substantially informs about\nits authorial writing style. Sentence representation learning has been widely\nexplored in recent years and it has been shown that it improves the\ngeneralization of different downstream tasks across many domains. Even though\nutilizing probing methods in several studies suggests that these learned\ncontextual representations implicitly encode some amount of syntax, explicit\nsyntactic information further improves the performance of deep neural models in\nthe domain of authorship attribution. These observations have motivated us to\ninvestigate the explicit representation learning of syntactic structure of\nsentences. In this paper, we propose a self-supervised framework for learning\nstructural representations of sentences. The self-supervised network contains\ntwo components; a lexical sub-network and a syntactic sub-network which take\nthe sequence of words and their corresponding structural labels as the input,\nrespectively. Due to the n-to-1 mapping of words to their structural labels,\neach word will be embedded into a vector representation which mainly carries\nstructural information. We evaluate the learned structural representations of\nsentences using different probing tasks, and subsequently utilize them in the\nauthorship attribution task. Our experimental results indicate that the\nstructural embeddings significantly improve the classification tasks when\nconcatenated with the existing pre-trained word embeddings.\n
Paper
References (60)
Scroll for more · 38 remaining