Automatic Speech Recognition (ASR) is an imperfect process that results in\ncertain mismatches in ASR output text when compared to plain written text or\ntranscriptions. When plain text data is to be used to train systems for spoken\nlanguage understanding or ASR, a proven strategy to reduce said mismatch and\nprevent degradations, is to hallucinate what the ASR outputs would be given a\ngold transcription. Prior work in this domain has focused on modeling errors at\nthe phonetic level, while using a lexicon to convert the phones to words,\nusually accompanied by an FST Language model. We present novel end-to-end\nmodels to directly predict hallucinated ASR word sequence outputs, conditioning\non an input word sequence as well as a corresponding phoneme sequence. This\nimproves prior published results for recall of errors from an in-domain ASR\nsystem's transcription of unseen data, as well as an out-of-domain ASR system's\ntranscriptions of audio from an unrelated task, while additionally exploring an\nin-between scenario when limited characterization data from the test ASR system\nis obtainable. To verify the extrinsic validity of the method, we also use our\nhallucinated ASR errors to augment training for a spoken question classifier,\nfinding that they enable robustness to real ASR errors in a downstream task,\nwhen scarce or even zero task-specific audio was available at train-time.\n
Paper
References (33)
Scroll for more · 21 remaining