中文

Aksharantar:面向下一个十亿用户的开放印度语言转写数据集与模型

计算与语言 2023-10-27 v2

摘要

由于多种文字的使用以及罗马化输入的广泛使用,转写在印度语言语境下非常重要。然而,公开可用的训练与评估集很少。我们介绍 Aksharantar,通过从单语与平行语料库挖掘以及收集人类标注者数据创建的、最大的公开可用印度语言转写数据集。该数据集包含来自 3 个语系、使用 12 种文字的 21 种印度语言的 2600 万转写对。Aksharantar 比现有数据集大 21 倍,并且是首个公开可用的涵盖 7 种语言与 1 个语系的数据集。我们还引入了 Aksharantar 测试集,包含跨越 19 种语言的 103k 词对,能够对转写模型在原生词、外来词、高频词与罕见词上进行细粒度分析。利用训练集,我们训练了 IndicXlit,一种多语言转写模型,在 Dakshina 测试集上准确率提升 15%,并在本工作引入的 Aksharantar 测试集上建立了强基线。模型、挖掘脚本、转写指南与数据集均在 https://github.com/AI4Bharat/IndicXlit 下以开源许可提供。我们希望这些大规模开放资源的可用性能激发印度语言转写及下游应用的创新。我们希望这些大规模开放资源的可用性能激发印度语言转写及下游应用的创新。

关键词

引用

@article{arxiv.2205.03018,
  title  = {Aksharantar: Open Indic-language Transliteration datasets and models for the Next Billion Users},
  author = {Yash Madhani and Sushane Parthan and Priyanka Bedekar and Gokul NC and Ruchi Khapra and Anoop Kunchukuttan and Pratyush Kumar and Mitesh M. Khapra},
  journal= {arXiv preprint arXiv:2205.03018},
  year   = {2023}
}

备注

This manuscript is an extended version of the paper accepted to EMNLP Findings 2023. You can find the EMNLP Findings version at https://anoopkunchukuttan.gitlab.io/publications/emnlp_findings_2023_aksharantar.pdf