Recognition errors are common in human communication. Similar errors often\nlead to unwanted behaviour in dialogue systems or virtual assistants. In human\ncommunication, we can recover from them by repeating misrecognized words or\nphrases; however in human-machine communication this recovery mechanism is not\navailable. In this paper, we attempt to bridge this gap and present a system\nthat allows a user to correct speech recognition errors in a virtual assistant\nby repeating misunderstood words. When a user repeats part of the phrase the\nsystem rewrites the original query to incorporate the correction. This rewrite\nallows the virtual assistant to understand the original query successfully. We\npresent an end-to-end 2-step attention pointer network that can generate the\nthe rewritten query by merging together the incorrectly understood utterance\nwith the correction follow-up. We evaluate the model on data collected for this\ntask and compare the proposed model to a rule-based baseline and a standard\npointer network. We show that rewriting the original query is an effective way\nto handle repetition-based recovery and that the proposed model outperforms the\nrule based baseline, reducing Word Error Rate by 19% relative at 2% False Alarm\nRate on annotated data.\n
Paper
References (23)
Scroll for more · 11 remaining