2019/01/11 by J. Edward Hu, Jingyue Hu, Hu, J. Edward +6 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.1901.03644
To be presented at AAAI 2019. 8 pages
arxiv created 2019/01/11 · openalex publication_date 2019/01/11 · arxiv updated 2019/01/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We present ParaBank, a large-scale English paraphrase dataset that surpasses prior work in both quantity and quality. Following the approach of ParaNMT, we train a Czech-English neural machine translation (NMT) system to generate novel paraphrases of English reference sentences. By adding lexical constraints to the NMT decoding procedure, however, we are able to produce multiple high-quality sentential paraphrases per source sentence, yielding an English paraphrase resource with more than 4 billion generated tokens and exhibiting greater lexical diversity. Using human judgments, we also demonstrate that ParaBank's paraphrases improve over ParaNMT on both semantic similarity and fluency. Finally, we use ParaBank to train a monolingual NMT model with the same support for lexically-constrained decoding for sentence rewriting tasks.