中文

关于 Jopara 情感分析的后勤困难与发现

计算与语言 2021-05-12 v2 机器学习

摘要

本文解决 Jopara(瓜拉尼语与西班牙语之间的语码转换语言)的情感分析问题。我们首先收集了一个以瓜拉尼语为主的推特语料库,并讨论即便对于情感分析这类相对易于标注的任务,也难以找到高质量数据。然后,我们训练了一系列神经模型,包括预训练语言模型,并探究在此低资源设定下它们是否优于传统机器学习模型。Transformer 架构取得了最佳结果,尽管预训练时未考虑瓜拉尼语,但由于问题的低资源性质,传统机器学习模型表现接近。

关键词

引用

@article{arxiv.2105.02947,
  title  = {On the logistical difficulties and findings of Jopara Sentiment Analysis},
  author = {Marvin M. Agüero-Torales and David Vilares and Antonio G. López-Herrera},
  journal= {arXiv preprint arXiv:2105.02947},
  year   = {2021}
}

备注

Accepted in the CALCS 2021 (co-located with NAACL 2021) - Fifth Workshop on Computational Approaches to Linguistic Code Switching, to appear (June 2021)