vix.ing · top · new · best · stats

Learning Multilingual Word Representations using a Bag-of-Words Autoencoder

2014/01/08 by Stanislas Lauly, Lauly, Stanislas, Alex Boulanger +3 · 1 citation
Computer Science · Mathematics · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #cs.CL #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.1401.1803

This workshop paper was accepted on Octoble 30 2013 at the NIPS 2013 workshop on deep learning (https://sites.google.com/site/deeplearningworkshopnips2013/accepted-papers)

arxiv created 2014/01/08 · openalex publication_date 2014/01/08 · arxiv updated 2014/01/09 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28

Abstract

Recent work on learning multilingual word representations usually relies on the use of word-level alignements (e.g. infered with the help of GIZA++) between translated sentences, in order to align the word embeddings in different languages. In this workshop paper, we investigate an autoencoder model for learning multilingual word representations that does without such word-level alignements. The autoencoder is trained to reconstruct the bag-of-word representation of given sentence from an encoded representation extracted from its translation. We evaluate our approach on a multilingual document classification task, where labeled data is available only for one language (e.g. English) while classification must be performed in a different language (e.g. French). In our experiments, we observe that our method compares favorably with a previously proposed method that exploits word-level alignments to learn word representations.

Citations

Cited by

Related