Summary Reinforcement learning from human feedback (RLHF) has shown strong potential for aligning language models, but its role in task‐oriented dialogue (TOD) remains unclear. In TOD, models are typically trained with local turn‐level supervision, while system behavior is evaluated through broader interaction‐level properties. This mismatch becomes more challenging in online settings, where explicit dialogue‐level rewards and human preference annotations are unavailable. In this work, we study whether RLHF can be usefully applied to TOD under this limitation. We consider two task‐annotation regimes, partially annotated and fully annotated TOD data, and construct pseudo‐preference pairs using empirical ranking heuristics motivated by prior work on synthetic feedback and model‐based ranking signals. We then train reward models on the constructed pairs and optimize dialogue policies with Preference policy optimization (PPO) using simulator‐generated online trajectories. Experiments on MultiWOZ 2.1 show that the proposed RLHF approach consistently improves corpus‐based evaluation over supervised baselines, while simulator‐based effects remain mixed and backbone‐dependent.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex