Reinforcement learning (RL) is frequently used to increase performance in\ntext generation tasks, including machine translation (MT), notably through the\nuse of Minimum Risk Training (MRT) and Generative Adversarial Networks (GAN).\nHowever, little is known about what and how these methods learn in the context\nof MT. We prove that one of the most common RL methods for MT does not optimize\nthe expected reward, as well as show that other methods take an infeasibly long\ntime to converge. In fact, our results suggest that RL practices in MT are\nlikely to improve performance only where the pre-trained parameters are already\nclose to yielding the correct translation. Our findings further suggest that\nobserved gains may be due to effects unrelated to the training signal, but\nrather from changes in the shape of the distribution curve.\n
Paper
References (48)
Scroll for more · 36 remaining