We introduce Sentence-level Language Modeling, a new pre-training objective\nfor learning a discourse language representation in a fully self-supervised\nmanner. Recent pre-training methods in NLP focus on learning either bottom or\ntop-level language representations: contextualized word representations derived\nfrom language model objectives at one extreme and a whole sequence\nrepresentation learned by order classification of two given textual segments at\nthe other. However, these models are not directly encouraged to capture\nrepresentations of intermediate-size structures that exist in natural languages\nsuch as sentences and the relationships among them. To that end, we propose a\nnew approach to encourage learning of a contextualized sentence-level\nrepresentation by shuffling the sequence of input sentences and training a\nhierarchical transformer model to reconstruct the original ordering. Through\nexperiments on downstream tasks such as GLUE, SQuAD, and DiscoEval, we show\nthat this feature of our model improves the performance of the original BERT by\nlarge margins.\n
Paper
References (33)
Scroll for more · 21 remaining