2019/07/05 by Chinnadhurai Sankar, Sujith Ravi, Sankar, Chinnadhurai +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.1907.02848
SIGDIAL 2019 - BEST PAPER AWARD
openalex publication_date 2019/07/05 · arxiv created 2019/09/14 · arxiv updated 2019/09/17 · openalex created_date 2022/07/28 · openalex updated_date 2026/07/28
Open domain dialog systems face the challenge of being repetitive and producing generic responses. In this paper, we demonstrate that by conditioning the response generation on interpretable discrete dialog attributes and composed attributes, it helps improve the model perplexity and results in diverse and interesting non-redundant responses. We propose to formulate the dialog attribute prediction as a reinforcement learning (RL) problem and use policy gradients methods to optimize utterance generation using long-term rewards. Unlike existing RL approaches which formulate the token prediction as a policy, our method reduces the complexity of the policy optimization by limiting the action space to dialog attributes, thereby making the policy optimization more practical and sample efficient. We demonstrate this with experimental and human evaluations.