中文

面向低资源语言的基于锚点的双语词嵌入

计算与语言 2021-07-28 v2

摘要

对于拥有大量无标注文本的语言,可以构建高质量的 monolingual word embeddings (MWEs,单语词嵌入)。仅需几千个词翻译对即可将 MWEs 对齐到双语空间。对于低资源语言,以单语方式训练 MWEs 会导致其质量较差,进而双语词嵌入 (BWEs) 的质量也较差。本文提出一种构建 BWEs 的新方法,该方法以高资源源语言的向量空间作为训练低资源目标语言嵌入空间的起点。通过将源语言向量用作锚点,向量空间在训练过程中自动对齐。我们在 English-German、English-Hiligaynon 和 English-Macedonian 上进行了实验。我们表明,我们的方法不仅提升了 BWEs 与双语词典归纳性能,还提升了以单语词相似性衡量的目标语言 MWE 质量。

关键词

引用

@article{arxiv.2010.12627,
  title  = {Anchor-based Bilingual Word Embeddings for Low-Resource Languages},
  author = {Tobias Eder and Viktor Hangya and Alexander Fraser},
  journal= {arXiv preprint arXiv:2010.12627},
  year   = {2021}
}

备注

The Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing