In recent years, Vietnam witnesses the mass development of social network users on different social platforms such as Facebook, Youtube, Instagram, and Tiktok. On social medias, hate speech has become a critical problem for social network users. To solve this problem, we introduce the ViHSD - a human-annotated dataset for automatically detecting hate speech on the social network. This dataset contains over 30,000 comments, each comment in the dataset has one of three labels: CLEAN, OFFENSIVE, or HATE. Besides, we introduce the data creation process for annotating and evaluating the quality of the dataset. Finally, we evaluated the dataset by deep learning models and transformer models.
@article{arxiv.2103.11528,
title = {A Large-scale Dataset for Hate Speech Detection on Vietnamese Social Media Texts},
author = {Son T. Luu and Kiet Van Nguyen and Ngan Luu-Thuy Nguyen},
journal= {arXiv preprint arXiv:2103.11528},
year = {2021}
}
Comments
IEA/AIE 2021: Advances and Trends in Artificial Intelligence. Artificial Intelligence Practices, pp 415-426