2019/09/09 by Lyan Verwimp, Verwimp, Lyan, Jerome R. Bellegarda +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1909.04130
openalex publication_date 2019/09/09 · openalex created_date 2022/07/19 · openalex updated_date 2026/07/28
Natural language processing (NLP) tasks tend to suffer from a paucity of\nsuitably annotated training data, hence the recent success of transfer learning\nacross a wide variety of them. The typical recipe involves: (i) training a\ndeep, possibly bidirectional, neural network with an objective related to\nlanguage modeling, for which training data is plentiful; and (ii) using the\ntrained network to derive contextual representations that are far richer than\nstandard linear word embeddings such as word2vec, and thus result in important\ngains. In this work, we wonder whether the opposite perspective is also true:\ncan contextual representations trained for different NLP tasks improve language\nmodeling itself? Since language models (LMs) are predominantly locally\noptimized, other NLP tasks may help them make better predictions based on the\nentire semantic fabric of a document. We test the performance of several types\nof pre-trained embeddings in neural LMs, and we investigate whether it is\npossible to make the LM more aware of global semantic information through\nembeddings pre-trained with a domain classification model. Initial experiments\nsuggest that as long as the proper objective criterion is used during training,\npre-trained embeddings are likely to be beneficial for neural language\nmodeling.\n