A Joint Learning Approach based on Self-Distillation for Keyphrase Extraction from Scientific Documents

Keyphrase extraction is the task of extracting a small set of phrases that\nbest describe a document. Most existing benchmark datasets for the task\ntypically have limited numbers of annotated documents, making it challenging to\ntrain increasingly complex neural networks. In contrast, digital libraries\nstore millions of scientific articles online, covering a wide range of topics.\nWhile a significant portion of these articles contain keyphrases provided by\ntheir authors, most other articles lack such kind of annotations. Therefore, to\neffectively utilize these large amounts of unlabeled articles, we propose a\nsimple and efficient joint learning approach based on the idea of\nself-distillation. Experimental results show that our approach consistently\nimproves the performance of baseline models for keyphrase extraction.\nFurthermore, our best models outperform previous methods for the task,\nachieving new state-of-the-art results on two public benchmarks: Inspec and\nSemEval-2017.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC