In recent studies, it has been shown that Multilingual language models underperform their monolingual counterparts. It is also a well-known fact that training and maintaining monolingual models for each language is a costly and time-consuming process. Roman Urdu is a resource-starved language used popularly on social media platforms and chat apps. In this research, we propose a novel dataset of scraped tweets containing 54M tokens and 3M sentences. Additionally, we also propose RUBERT a bilingual Roman Urdu model created by additional pretraining of English BERT. We compare its performance with a monolingual Roman Urdu BERT trained from scratch and a multilingual Roman Urdu BERT created by additional pretraining of Multilingual BERT. We show through our experiments that additional pretraining of the English BERT produces the most notable performance improvement.
@article{arxiv.2102.11278,
title = {RUBERT: A Bilingual Roman Urdu BERT Using Cross Lingual Transfer Learning},
author = {Usama Khalid and Mirza Omer Beg and Muhammad Umair Arshad},
journal= {arXiv preprint arXiv:2102.11278},
year = {2021}
}
Comments
arXiv admin note: substantial text overlap with arXiv:2102.10958