To Point or Not to Point: Understanding How Abstractive Summarizers Paraphrase Text

Abstractive neural summarization models have seen great improvements in\nrecent years, as shown by ROUGE scores of the generated summaries. But despite\nthese improved metrics, there is limited understanding of the strategies\ndifferent models employ, and how those strategies relate their understanding of\nlanguage. To understand this better, we run several experiments to characterize\nhow one popular abstractive model, the pointer-generator model of See et al.\n(2017), uses its explicit copy/generation switch to control its level of\nabstraction (generation) vs extraction (copying). On an extractive-biased\ndataset, the model utilizes syntactic boundaries to truncate sentences that are\notherwise often copied verbatim. When we modify the copy/generation switch and\nforce the model to generate, only simple paraphrasing abilities are revealed\nalongside factual inaccuracies and hallucinations. On an abstractive-biased\ndataset, the model copies infrequently but shows similarly limited abstractive\nabilities. In line with previous research, these results suggest that\nabstractive summarization models lack the semantic understanding necessary to\ngenerate paraphrases that are both abstractive and faithful to the source\ndocument.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC