Despite recent advances in neural text generation, encoding the rich\ndiversity in human language remains elusive. We argue that the sub-optimal text\ngeneration is mainly attributable to the imbalanced token distribution, which\nparticularly misdirects the learning model when trained with the\nmaximum-likelihood objective. As a simple yet effective remedy, we propose two\nnovel methods, F^2-Softmax and MefMax, for a balanced training even with the\nskewed frequency distribution. MefMax assigns tokens uniquely to frequency\nclasses, trying to group tokens with similar frequencies and equalize frequency\nmass between the classes. F^2-Softmax then decomposes a probability\ndistribution of the target token into a product of two conditional\nprobabilities of (i) frequency class, and (ii) token from the target frequency\nclass. Models learn more uniform probability distributions because they are\nconfined to subsets of vocabularies. Significant performance gains on seven\nrelevant metrics suggest the supremacy of our approach in improving not only\nthe diversity but also the quality of generated texts.\n
Paper
References (45)
Scroll for more · 33 remaining