基于语言向量的连续多语言性
计算与语言
2017-03-21 v2
摘要
现有的大多数多语言自然语言处理(NLP)模型将语言视为离散类别,并针对某一种或另一种语言进行预测。相比之下,我们提出使用语言的连续向量表示。我们证明这些表示可以通过基于字符的神经语言模型高效学习,并用于改善对训练中未见过的语言变体的推断。在基于 990 种不同语言的 1303 个《圣经》译本的实验中,我们实证探索了多语言模型的能力,并表明语言向量能够捕捉语言之间的谱系关系。
引用
@article{arxiv.1612.07486,
title = {Continuous multilinguality with language vectors},
author = {Robert Östling and Jörg Tiedemann},
journal= {arXiv preprint arXiv:1612.07486},
year = {2017}
}
备注
In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL), Valencia, Spain, April, 2017