Instant One-Shot Word-Learning for Context-Specific Neural Sequence-to-Sequence Speech Recognition

Neural sequence-to-sequence systems deliver state-of-the-art performance for\nautomatic speech recognition (ASR). When using appropriate modeling units,\ne.g., byte-pair encoded characters, these systems are in principal open\nvocabulary systems. In practice, however, they often fail to recognize words\nnot seen during training, e.g., named entities, numbers or technical terms. To\nalleviate this problem we supplement an end-to-end ASR system with a\nword/phrase memory and a mechanism to access this memory to recognize the words\nand phrases correctly. After the training of the ASR system, and when it has\nalready been deployed, a relevant word can be added or subtracted instantly\nwithout the need for further training. In this paper we demonstrate that\nthrough this mechanism our system is able to recognize more than 85% of newly\nadded words that it previously failed to recognize compared to a strong\nbaseline.\n

Paper

References (27)

Scroll for more · 15 remaining

Similar papers

© 2026 NYSGPT2525 LLC