中文

用于国家级别方言识别的多方言阿拉伯语 BERT

计算与语言 2020-07-14 v1 机器学习

摘要

阿拉伯语方言识别由于语言本身的一些固有特性而是一个复杂的问题。在本文中,我们介绍了我们的参赛团队 Mawdoo3 AI 在赢得 Nuanced Arabic Dialect Identification(NADI)共享任务子任务 1 的过程中所进行的实验和开发的模型。该方言识别子任务提供了覆盖所有 21 个阿拉伯国家的 21,000 条国家级别标注推文。竞赛组织者还提供了来自同一领域的 10M 条无标注推文语料库供可选使用。我们的获胜方案本身以我们预训练 BERT 模型的不同训练迭代的集成形式呈现,在该子任务上取得了 26.78% 的微平均 F1-score。我们公开发布了获胜方案中预训练语言模型组件,命名为 Multi-dialect-Arabic-BERT 模型,供任何感兴趣的研究者使用。

关键词

引用

@article{arxiv.2007.05612,
  title  = {Multi-Dialect Arabic BERT for Country-Level Dialect Identification},
  author = {Bashar Talafha and Mohammad Ali and Muhy Eddin Za'ter and Haitham Seelawi and Ibraheem Tuffaha and Mostafa Samir and Wael Farhan and Hussein T. Al-Natsheh},
  journal= {arXiv preprint arXiv:2007.05612},
  year   = {2020}
}

备注

Accepted at the Fifth Arabic Natural Language Processing Workshop (WANLP2020) co-located with the 28th International Conference on Computational Linguistics (COLING'2020), Barcelona, Spain, 12 Dec. 2020