We present a data-driven, end-to-end approach to transaction-based dialog\nsystems that performs at near-human levels in terms of verbal response quality\nand factual grounding accuracy. We show that two essential components of the\nsystem produce these results: a sufficiently large and diverse, in-domain\nlabeled dataset, and a neural network-based, pre-trained model that generates\nboth verbal responses and API call predictions. In terms of data, we introduce\nTicketTalk, a movie ticketing dialog dataset with 23,789 annotated\nconversations. The movie ticketing conversations range from completely\nopen-ended and unrestricted to more structured, both in terms of their\nknowledge base, discourse features, and number of turns. In qualitative human\nevaluations, model-generated responses trained on just 10,000 TicketTalk\ndialogs were rated to "make sense" 86.5 percent of the time, almost the same as\nhuman responses in the same contexts. Our simple, API-focused annotation schema\nresults in a much easier labeling task making it faster and more cost\neffective. It is also the key component for being able to predict API calls\naccurately. We handle factual grounding by incorporating API calls in the\ntraining data, allowing our model to learn which actions to take and when.\nTrained on the same 10,000-dialog set, the model's API call predictions were\nrated to be correct 93.9 percent of the time in our evaluations, surpassing the\nratings for the corresponding human labels. We show how API prediction and\nresponse generation scores improve as the dataset size incrementally increases\nfrom 5000 to 21,000 dialogs. Our analysis also clearly illustrates the benefits\nof pre-training. We are publicly releasing the TicketTalk dataset with this\npaper to facilitate future work on transaction-based dialogs.\n
Paper
References (33)
Scroll for more · 21 remaining