Likelihood training and maximization-based decoding result in dull and\nrepetitive generated texts even when using powerful language models (Holtzman\net al., 2019). Adding a loss function for regularization was shown to improve\ntext generation output by helping avoid unwanted properties, such as\ncontradiction or repetition (Li at al., 2020). In this work, we propose\nfine-tuning a language model by using policy gradient reinforcement learning,\ndirectly optimizing for better generation. We apply this approach to minimizing\nrepetition in generated text, and show that, when combined with unlikelihood\ntraining (Welleck et al., 2020), our method further reduces repetition without\nimpacting the language model quality. We also evaluate other methods for\nimproving generation at training and decoding time, and compare them using\nvarious metrics aimed at control for better text generation output.\n
Paper
References (20)
Scroll for more · 8 remaining