Learning to Count Words in Fluent Speech enables Online Speech Recognition

Sequence to Sequence models, in particular the Transformer, achieve state of\nthe art results in Automatic Speech Recognition. Practical usage is however\nlimited to cases where full utterance latency is acceptable. In this work we\nintroduce Taris, a Transformer-based online speech recognition system aided by\nan auxiliary task of incremental word counting. We use the cumulative word sum\nto dynamically segment speech and enable its eager decoding into words.\nExperiments performed on the LRS2, LibriSpeech, and Aishell-1 datasets of\nEnglish and Mandarin speech show that the online system performs comparable\nwith the offline one when having a dynamic algorithmic delay of 5 segments.\nFurthermore, we show that the estimated segment length distribution resembles\nthe word length distribution obtained with forced alignment, although our\nsystem does not require an exact segment-to-word equivalence. Taris introduces\na negligible overhead compared to a standard Transformer, while the local\nrelationship modelling between inputs and outputs grants invariance to sequence\nlength by design.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC