2020/08/17 by Michael Sejr Schlichtkrull, Michael Schlichtkrull, Weiwei Cheng +2
Computer Science · Mathematics · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (stat.ML) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL #stat.ML
paper · pdf · doi:10.48550/arxiv.2008.07291
arxiv created 2020/08/17 · openalex publication_date 2020/08/17 · arxiv updated 2020/08/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Generating diverse and relevant questions over text is a task with widespread applications. We argue that commonly-used evaluation metrics such as BLEU and METEOR are not suitable for this task due to the inherent diversity of reference questions, and propose a scheme for extending conventional metrics to reflect diversity. We furthermore propose a variational encoder-decoder model for this task. We show through automatic and human evaluation that our variational model improves diversity without loss of quality, and demonstrate how our evaluation scheme reflects this improvement.