English

AraMUS: Pushing the Limits of Data and Model Scale for Arabic Natural Language Processing

Computation and Language 2023-06-13 v1

Abstract

Developing monolingual large Pre-trained Language Models (PLMs) is shown to be very successful in handling different tasks in Natural Language Processing (NLP). In this work, we present AraMUS, the largest Arabic PLM with 11B parameters trained on 529GB of high-quality Arabic textual data. AraMUS achieves state-of-the-art performances on a diverse set of Arabic classification and generative tasks. Moreover, AraMUS shows impressive few-shot learning abilities compared with the best existing Arabic PLMs.

Keywords

Cite

@article{arxiv.2306.06800,
  title  = {AraMUS: Pushing the Limits of Data and Model Scale for Arabic Natural Language Processing},
  author = {Asaad Alghamdi and Xinyu Duan and Wei Jiang and Zhenhai Wang and Yimeng Wu and Qingrong Xia and Zhefeng Wang and Yi Zheng and Mehdi Rezagholizadeh and Baoxing Huai and Peilun Cheng and Abbas Ghaddar},
  journal= {arXiv preprint arXiv:2306.06800},
  year   = {2023}
}