English

BERTi\'c -- The Transformer Language Model for Bosnian, Croatian, Montenegrin and Serbian

Computation and Language 2021-04-20 v1

Abstract

In this paper we describe a transformer model pre-trained on 8 billion tokens of crawled text from the Croatian, Bosnian, Serbian and Montenegrin web domains. We evaluate the transformer model on the tasks of part-of-speech tagging, named-entity-recognition, geo-location prediction and commonsense causal reasoning, showing improvements on all tasks over state-of-the-art models. For commonsense reasoning evaluation, we introduce COPA-HR -- a translation of the Choice of Plausible Alternatives (COPA) dataset into Croatian. The BERTi\'c model is made available for free usage and further task-specific fine-tuning through HuggingFace.

Keywords

Cite

@article{arxiv.2104.09243,
  title  = {BERTi\'c -- The Transformer Language Model for Bosnian, Croatian, Montenegrin and Serbian},
  author = {Nikola Ljubešić and Davor Lauc},
  journal= {arXiv preprint arXiv:2104.09243},
  year   = {2021}
}