2015/10/11 by Jiwei Li, Michel Galley, Li, Jiwei +7 · 167 citations
Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.1510.03055
In. Proc of NAACL 2016
arxiv created 2016/06/10 · arxiv updated 2016/06/14
Sequence-to-sequence neural network models for generation of conversational responses tend to generate safe, commonplace responses (e.g., "I don't know") regardless of the input. We suggest that the traditional objective function, i.e., the likelihood of output (response) given input (message) is unsuited to response generation tasks. Instead we propose using Maximum Mutual Information (MMI) as the objective function in neural models. Experimental results demonstrate that the proposed MMI models produce more diverse, interesting, and appropriate responses, yielding substantive gains in BLEU scores on two conversational datasets and in human evaluations.