2021/04/06 by Hassan S. Shavarani, Shavarani, Hassan S., Anoop Sarkar +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2104.02831
openalex publication_date 2021/04/06 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28
Adding linguistic information (syntax or semantics) to neural machine\ntranslation (NMT) has mostly focused on using point estimates from pre-trained\nmodels. Directly using the capacity of massive pre-trained contextual word\nembedding models such as BERT (Devlin et al., 2019) has been marginally useful\nin NMT because effective fine-tuning is difficult to obtain for NMT without\nmaking training brittle and unreliable. We augment NMT by extracting dense\nfine-tuned vector-based linguistic information from BERT instead of using point\nestimates. Experimental results show that our method of incorporating\nlinguistic information helps NMT to generalize better in a variety of training\ncontexts and is no more difficult to train than conventional Transformer-based\nNMT.\n