English

On the logistical difficulties and findings of Jopara Sentiment Analysis

Computation and Language 2021-05-12 v2 Machine Learning

Abstract

This paper addresses the problem of sentiment analysis for Jopara, a code-switching language between Guarani and Spanish. We first collect a corpus of Guarani-dominant tweets and discuss on the difficulties of finding quality data for even relatively easy-to-annotate tasks, such as sentiment analysis. Then, we train a set of neural models, including pre-trained language models, and explore whether they perform better than traditional machine learning ones in this low-resource setup. Transformer architectures obtain the best results, despite not considering Guarani during pre-training, but traditional machine learning models perform close due to the low-resource nature of the problem.

Keywords

Cite

@article{arxiv.2105.02947,
  title  = {On the logistical difficulties and findings of Jopara Sentiment Analysis},
  author = {Marvin M. Agüero-Torales and David Vilares and Antonio G. López-Herrera},
  journal= {arXiv preprint arXiv:2105.02947},
  year   = {2021}
}

Comments

Accepted in the CALCS 2021 (co-located with NAACL 2021) - Fifth Workshop on Computational Approaches to Linguistic Code Switching, to appear (June 2021)