Are Neural Open-Domain Dialog Systems Robust to Speech Recognition Errors in the Dialog History? An Empirical Study
Large end-to-end neural open-domain chatbots are becoming increasingly\npopular. However, research on building such chatbots has typically assumed that\nthe user input is written in nature and it is not clear whether these chatbots\nwould seamlessly integrate with automatic speech recognition (ASR) models to\nserve the speech modality. We aim to bring attention to this important question\nby empirically studying the effects of various types of synthetic and actual\nASR hypotheses in the dialog history on TransferTransfo, a state-of-the-art\nGenerative Pre-trained Transformer (GPT) based neural open-domain dialog system\nfrom the NeurIPS ConvAI2 challenge. We observe that TransferTransfo trained on\nwritten data is very sensitive to such hypotheses introduced to the dialog\nhistory during inference time. As a baseline mitigation strategy, we introduce\nsynthetic ASR hypotheses to the dialog history during training and observe\nmarginal improvements, demonstrating the need for further research into\ntechniques to make end-to-end open-domain chatbots fully speech-robust. To the\nbest of our knowledge, this is the first study to evaluate the effects of\nsynthetic and actual ASR hypotheses on a state-of-the-art neural open-domain\ndialog system and we hope it promotes speech-robustness as an evaluation\ncriterion in open-domain dialog.\n