Over its three decade history, speech translation has experienced several\nshifts in its primary research themes; moving from loosely coupled cascades of\nspeech recognition and machine translation, to exploring questions of tight\ncoupling, and finally to end-to-end models that have recently attracted much\nattention. This paper provides a brief survey of these developments, along with\na discussion of the main challenges of traditional approaches which stem from\ncommitting to intermediate representations from the speech recognizer, and from\ntraining cascaded models separately towards different objectives.\n Recent end-to-end modeling techniques promise a principled way of overcoming\nthese issues by allowing joint training of all model components and removing\nthe need for explicit intermediate representations. However, a closer look\nreveals that many end-to-end models fall short of solving these issues, due to\ncompromises made to address data scarcity. This paper provides a unifying\ncategorization and nomenclature that covers both traditional and recent\napproaches and that may help researchers by highlighting both trade-offs and\nopen research questions.\n
Paper
References (95)
Scroll for more · 38 remaining