Generalized Entropy Regularization or: There's Nothing Special about Label Smoothing

Prior work has explored directly regularizing the output distributions of\nprobabilistic models to alleviate peaky (i.e. over-confident) predictions, a\ncommon sign of overfitting. This class of techniques, of which label smoothing\nis one, has a connection to entropy regularization. Despite the consistent\nsuccess of label smoothing across architectures and data sets in language\ngeneration tasks, two problems remain open: (1) there is little understanding\nof the underlying effects entropy regularizers have on models, and (2) the full\nspace of entropy regularization techniques is largely unexplored. We introduce\na parametric family of entropy regularizers, which includes label smoothing as\na special case, and use it to gain a better understanding of the relationship\nbetween the entropy of a model and its performance on language generation\ntasks. We also find that variance in model performance can be explained largely\nby the resulting entropy of the model. Lastly, we find that label smoothing\nprovably does not allow for sparsity in an output distribution, an undesirable\nproperty for language generation models, and therefore advise the use of other\nentropy regularization methods in its place.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC