vix.ing · top · new · best · stats · spec

ParaBank: Monolingual Bitext Generation and Sentential Paraphrasing via\n Lexically-constrained Neural Machine Translation

2019/01/11 by Jingyue Hu, Hu, J. Edward, Rachel Rudinger +5 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1901.03644

openalex publication_date 2019/01/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We present ParaBank, a large-scale English paraphrase dataset that surpasses\nprior work in both quantity and quality. Following the approach of ParaNMT, we\ntrain a Czech-English neural machine translation (NMT) system to generate novel\nparaphrases of English reference sentences. By adding lexical constraints to\nthe NMT decoding procedure, however, we are able to produce multiple\nhigh-quality sentential paraphrases per source sentence, yielding an English\nparaphrase resource with more than 4 billion generated tokens and exhibiting\ngreater lexical diversity. Using human judgments, we also demonstrate that\nParaBank's paraphrases improve over ParaNMT on both semantic similarity and\nfluency. Finally, we use ParaBank to train a monolingual NMT model with the\nsame support for lexically-constrained decoding for sentence rewriting tasks.\n

Cited by

Related