2020/05/16 by Chao Xiong, Xiong, Chao, Che Liu +7
Computer Science · Mathematics · #AI in Service Interactions #Artificial intelligence #Computation and Language (cs.CL) #Computer science #Context (archaeology) #Domain (mathematical analysis) #FOS: Computer and information sciences #Image (mathematics) #Information retrieval #Machine Learning (cs.LG) #Matching (statistics) #Mathematics #Natural Language Processing Techniques #Natural language processing #Open domain #Question answering #Selection (genetic algorithm) #Semantic matching #Semantic similarity #Sentence #Similarity (geometry) #Topic Modeling #Word (group theory) #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2005.07923
published in arXiv (Cornell University) (Cornell University) · 10 pages, 4 figures
arxiv created 2020/05/16 · openalex publication_date 2020/05/16 · arxiv updated 2020/05/19 · openalex created_date 2020/05/21 · openalex updated_date 2026/08/08
Recently, open domain multi-turn chatbots have attracted much interest from lots of researchers in both academia and industry. The dominant retrieval-based methods use context-response matching mechanisms for multi-turn response selection. Specifically, the state-of-the-art methods perform the context-response matching by word or segment similarity. However, these models lack a full exploitation of the sentence-level semantic information, and make simple mistakes that humans can easily avoid. In this work, we propose a matching network, called sequential sentence matching network (S2M), to use the sentence-level semantic information to address the problem. Firstly and most importantly, we find that by using the sentence-level semantic information, the network successfully addresses the problem and gets a significant improvement on matching, resulting in a state-of-the-art performance. Furthermore, we integrate the sentence matching we introduced here and the usual word similarity matching reported in the current literature, to match at different semantic levels. Experiments on three public data sets show that such integration further improves the model performance.