Reinforcement learning (RL) can enable task-oriented dialogue systems to\nsteer the conversation towards successful task completion. In an end-to-end\nsetting, a response can be constructed in a word-level sequential decision\nmaking process with the entire system vocabulary as action space. Policies\ntrained in such a fashion do not require expert-defined action spaces, but they\nhave to deal with large action spaces and long trajectories, making RL\nimpractical. Using the latent space of a variational model as action space\nalleviates this problem. However, current approaches use an uninformed prior\nfor training and optimize the latent distribution solely on the context. It is\ntherefore unclear whether the latent representation truly encodes the\ncharacteristics of different actions. In this paper, we explore three ways of\nleveraging an auxiliary task to shape the latent variable distribution: via\npre-training, to obtain an informed prior, and via multitask learning. We\nchoose response auto-encoding as the auxiliary task, as this captures the\ngenerative factors of dialogue responses while requiring low computational cost\nand neither additional data nor labels. Our approach yields a more\naction-characterized latent representations which support end-to-end dialogue\npolicy optimization and achieves state-of-the-art success rates. These results\nwarrant a more wide-spread use of RL in end-to-end dialogue models.\n
Paper
References (41)
Scroll for more · 29 remaining