Paraphrase Generation as Zero-Shot Multilingual Translation: Disentangling Semantic Similarity from Lexical and Syntactic Diversity
Recent work has shown that a multilingual neural machine translation (NMT)\nmodel can be used to judge how well a sentence paraphrases another sentence in\nthe same language (Thompson and Post, 2020); however, attempting to generate\nparaphrases from such a model using standard beam search produces trivial\ncopies or near copies. We introduce a simple paraphrase generation algorithm\nwhich discourages the production of n-grams that are present in the input. Our\napproach enables paraphrase generation in many languages from a single\nmultilingual NMT model. Furthermore, the amount of lexical diversity between\nthe input and output can be controlled at generation time. We conduct a human\nevaluation to compare our method to a paraphraser trained on the large English\nsynthetic paraphrase database ParaBank 2 (Hu et al., 2019c) and find that our\nmethod produces paraphrases that better preserve meaning and are more\ngramatical, for the same level of lexical diversity. Additional smaller human\nassessments demonstrate our approach also works in two non-English languages.\n
Paper
References (50)
Scroll for more · 38 remaining