2018/04/28 by Weiyan Shi, Zhou Yu, Shi, Weiyan +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1804.10731
openalex publication_date 2018/04/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
End-to-end learning framework is useful for building dialog systems for its simplicity in training and efficiency in model updating. However, current end-to-end approaches only consider user semantic inputs in learning and under-utilize other user information. Therefore, we propose to include user sentiment obtained through multimodal information (acoustic, dialogic and textual), in the end-to-end learning framework to make systems more user-adaptive and effective. We incorporated user sentiment information in both supervised and reinforcement learning settings. In both settings, adding sentiment information reduced the dialog length and improved the task success rate on a bus information search task. This work is the first attempt to incorporate multimodal user information in the adaptive end-to-end dialog system training framework and attained state-of-the-art performance.