中文

去偏多语言词嵌入:以三种印度语言为例

计算与语言 2021-07-23 v2

摘要

本文中,我们推进了当前最先进的单语言词嵌入去偏方法,使其能在多语言设定中良好泛化。我们考虑了量化偏见的多种方法以及单语言与多语言设定下的不同去偏途径。我们展示了偏见缓解方法在下游 NLP 应用中的重要性。我们提出的方法为 Hindi、Bengali 与 Telugu 三种印度语言及英语的多语言嵌入去偏确立了最先进性能。我们相信我们的工作将为构建无偏下游 NLP 应用开辟新机遇,此类应用固有地依赖于所用词嵌入的质量。

关键词

引用

@article{arxiv.2107.10181,
  title  = {Debiasing Multilingual Word Embeddings: A Case Study of Three Indian Languages},
  author = {Srijan Bansal and Vishal Garimella and Ayush Suhane and Animesh Mukherjee},
  journal= {arXiv preprint arXiv:2107.10181},
  year   = {2021}
}

备注

This work is accepted as a long paper in the proceedings of ACM HyperText 2021