We examine learning offensive content on Twitter with limited, imbalanced data. For the purpose, we investigate the utility of using various data enhancement methods with a host of classical ensemble classifiers. Among the 75 participating teams in SemEval-2019 sub-task B, our system ranks 6th (with 0.706 macro F1-score). For sub-task C, among the 65 participating teams, our system ranks 9th (with 0.587 macro F1-score).
@article{arxiv.1906.03692,
title = {UBC-NLP at SemEval-2019 Task 6:Ensemble Learning of Offensive Content With Enhanced Training Data},
author = {Arun Rajendran and Chiyu Zhang and Muhammad Abdul-Mageed},
journal= {arXiv preprint arXiv:1906.03692},
year = {2019}
}
Comments
7 pages, 2 figures, Proceedings of the 13th International Workshop on Semantic Evaluation (SemEval)