English

EENLP: Cross-lingual Eastern European NLP Index

Computation and Language 2022-05-12 v3 Artificial Intelligence Neural and Evolutionary Computing

Abstract

Motivated by the sparsity of NLP resources for Eastern European languages, we present a broad index of existing Eastern European language resources (90+ datasets and 45+ models) published as a github repository open for updates from the community. Furthermore, to support the evaluation of commonsense reasoning tasks, we provide hand-crafted cross-lingual datasets for five different semantic tasks (namely news categorization, paraphrase detection, Natural Language Inference (NLI) task, tweet sentiment detection, and news sentiment detection) for some of the Eastern European languages. We perform several experiments with the existing multilingual models on these datasets to define the performance baselines and compare them to the existing results for other languages.

Keywords

Cite

@article{arxiv.2108.02605,
  title  = {EENLP: Cross-lingual Eastern European NLP Index},
  author = {Alexey Tikhonov and Alex Malkhasov and Andrey Manoshin and George Dima and Réka Cserháti and Md. Sadek Hossain Asif and Matt Sárdi},
  journal= {arXiv preprint arXiv:2108.02605},
  year   = {2022}
}

Comments

Accepted for LREC 2022. 5 pages, 1 figure. Originally EEML 2021 project