English

Bilingual Learning of Multi-sense Embeddings with Discrete Autoencoders

Computation and Language 2016-03-31 v1 Machine Learning Machine Learning

Abstract

We present an approach to learning multi-sense word embeddings relying both on monolingual and bilingual information. Our model consists of an encoder, which uses monolingual and bilingual context (i.e. a parallel sentence) to choose a sense for a given word, and a decoder which predicts context words based on the chosen sense. The two components are estimated jointly. We observe that the word representations induced from bilingual data outperform the monolingual counterparts across a range of evaluation tasks, even though crosslingual information is not available at test time.

Keywords

Cite

@article{arxiv.1603.09128,
  title  = {Bilingual Learning of Multi-sense Embeddings with Discrete Autoencoders},
  author = {Simon Šuster and Ivan Titov and Gertjan van Noord},
  journal= {arXiv preprint arXiv:1603.09128},
  year   = {2016}
}

Comments

11 pages, to appear at NAACL 2016