Subword Regularization: An Analysis of Scalability and Generalization\n for End-to-End Automatic Speech Recognition

Subwords are the most widely used output units in end-to-end speech\nrecognition. They combine the best of two worlds by modeling the majority of\nfrequent words directly and at the same time allow open vocabulary speech\nrecognition by backing off to shorter units or characters to construct words\nunseen during training. However, mapping text to subwords is ambiguous and\noften multiple segmentation variants are possible. Yet, many systems are\ntrained using only the most likely segmentation. Recent research suggests that\nsampling subword segmentations during training acts as a regularizer for neural\nmachine translation and speech recognition models, leading to performance\nimprovements. In this work, we conduct a principled investigation on the\nregularizing effect of the subword segmentation sampling method for a streaming\nend-to-end speech recognition task. In particular, we evaluate the subword\nregularization contribution depending on the size of the training dataset. Our\nresults suggest that subword regularization provides a consistent improvement\nof (2-8%) relative word-error-rate reduction, even in a large-scale setting\nwith datasets up to a size of 20k hours. Further, we analyze the effect of\nsubword regularization on recognition of unseen words and its implications on\nbeam diversity.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC