vix.ing · top · new · best · stats · spec

Sample-Efficient Model-based Actor-Critic for an Interactive Dialogue\n Task

2020/04/28 by Katya Kudashkina, Valliappa Chockalingam, Kudashkina, Katya +5
Computer Science · Engineering · #Adaptive Dynamic Programming Control #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #Evacuation and Crowd Dynamics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multi-Agent Systems and Negotiation #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2004.13657

openalex publication_date 2020/04/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Human-computer interactive systems that rely on machine learning are becoming\nparamount to the lives of millions of people who use digital assistants on a\ndaily basis. Yet, further advances are limited by the availability of data and\nthe cost of acquiring new samples. One way to address this problem is by\nimproving the sample efficiency of current approaches. As a solution path, we\npresent a model-based reinforcement learning algorithm for an interactive\ndialogue task. We build on commonly used actor-critic methods, adding an\nenvironment model and planner that augments a learning agent to learn the model\nof the environment dynamics. Our results show that, on a simulation that mimics\nthe interactive task, our algorithm requires 70 times fewer samples, compared\nto the baseline of commonly used model-free algorithm, and demonstrates 2~times\nbetter performance asymptotically. Moreover, we introduce a novel contribution\nof computing a soft planner policy and further updating a model-free policy\nyielding a less computationally expensive model-free agent as good as the\nmodel-based one. This model-based architecture serves as a foundation that can\nbe extended to other human-computer interactive tasks allowing further advances\nin this direction.\n

Citations

Related