vix.ing · top · new · best · stats · spec

Variational Reward Estimator Bottleneck: Learning Robust Reward\n Estimator for Multi-Domain Task-Oriented Dialog

2020/05/30 by Jeiyoon Park, Chanhee Lee, Park, Jeiyoon +5
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Speech and dialogue systems #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2006.00417

openalex publication_date 2020/05/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Despite its notable success in adversarial learning approaches to\nmulti-domain task-oriented dialog system, training the dialog policy via\nadversarial inverse reinforcement learning often fails to balance the\nperformance of the policy generator and reward estimator. During optimization,\nthe reward estimator often overwhelms the policy generator and produces\nexcessively uninformative gradients. We proposes the Variational Reward\nestimator Bottleneck (VRB), which is an effective regularization method that\naims to constrain unproductive information flows between inputs and the reward\nestimator. The VRB focuses on capturing discriminative features, by exploiting\ninformation bottleneck on mutual information. Empirical results on a\nmulti-domain task-oriented dialog dataset demonstrate that the VRB\nsignificantly outperforms previous methods.\n

Citations

Related