Single headed attention based sequence-to-sequence model for state-of-the-art results on Switchboard
It is generally believed that direct sequence-to-sequence (seq2seq) speech\nrecognition models are competitive with hybrid models only when a large amount\nof data, at least a thousand hours, is available for training. In this paper,\nwe show that state-of-the-art recognition performance can be achieved on the\nSwitchboard-300 database using a single headed attention, LSTM based model.\nUsing a cross-utterance language model, our single-pass speaker independent\nsystem reaches 6.4% and 12.5% word error rate (WER) on the Switchboard and\nCallHome subsets of Hub5'00, without a pronunciation lexicon. While careful\nregularization and data augmentation are crucial in achieving this level of\nperformance, experiments on Switchboard-2000 show that nothing is more useful\nthan more data. Overall, the combination of various regularizations and a\nsimple but fairly large model results in a new state of the art, 4.7% and 7.8%\nWER on the Switchboard and CallHome sets, using SWB-2000 without any external\ndata resources.\n