BERT, a pre-trained Transformer model, has achieved ground-breaking performance on multiple NLP tasks. In this paper, we describe BERTSUM, a simple variant of BERT, for extractive summarization. Our system is the state of the art on the CNN/Dailymail dataset, outperforming the previous best-performed system by 1.65 on ROUGE-L. The codes to reproduce our results are available at this https URL
Paper
References (18)
12Layer NormalizationJimmy Ba, J. Kiros, Geoffrey E. Hinton2016 · arXiv.org · 13k citations In Library
Scroll for more · 6 remaining