Diluted Near-Optimal Expert Demonstrations for Guiding Dialogue Stochastic Policy Optimisation

A learning dialogue agent can infer its behaviour from interactions with the\nusers. These interactions can be taken from either human-to-human or\nhuman-machine conversations. However, human interactions are scarce and costly,\nmaking learning from few interactions essential. One solution to speedup the\nlearning process is to guide the agent's exploration with the help of an\nexpert. We present in this paper several imitation learning strategies for\ndialogue policy where the guiding expert is a near-optimal handcrafted policy.\nWe incorporate these strategies with state-of-the-art reinforcement learning\nmethods based on Q-learning and actor-critic. We notably propose a randomised\nexploration policy which allows for a seamless hybridisation of the learned\npolicy and the expert. Our experiments show that our hybridisation strategy\noutperforms several baselines, and that it can accelerate the learning when\nfacing real humans.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC