2018/04/29 by Taku Kudo, Kudo, Taku · 58 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1804.10959
openalex publication_date 2018/04/29 · openalex created_date 2022/08/20 · openalex updated_date 2026/07/28
Subword units are an effective way to alleviate the open vocabulary problems\nin neural machine translation (NMT). While sentences are usually converted into\nunique subword sequences, subword segmentation is potentially ambiguous and\nmultiple segmentations are possible even with the same vocabulary. The question\naddressed in this paper is whether it is possible to harness the segmentation\nambiguity as a noise to improve the robustness of NMT. We present a simple\nregularization method, subword regularization, which trains the model with\nmultiple subword segmentations probabilistically sampled during training. In\naddition, for better subword sampling, we propose a new subword segmentation\nalgorithm based on a unigram language model. We experiment with multiple\ncorpora and report consistent improvements especially on low resource and\nout-of-domain settings.\n