基于双向 LSTM 模型的越南语社交媒体文本仇恨言论检测
计算与语言
2019-11-12 v1 机器学习
摘要
在本文中,我们描述了参与 VLSP 2019 评测活动“社交媒体上的仇恨言论检测”共享任务的系统。我们获得了用于社交媒体评论或帖子的预标注数据集和未标注数据集。我们的任务是对评论/帖子进行预处理并构建机器学习模型以分类。在本报告中,我们使用双向长短期记忆(Bidirectional Long Short-Term Memory)构建模型,该模型可依据 Clean、Offensive、Hate 为社交媒体文本预测标签。凭借该系统,我们在 VLSP 2019 公开标准测试集上取得了 71.43% 的对比结果。
引用
@article{arxiv.1911.03648,
title = {Hate Speech Detection on Vietnamese Social Media Text using the Bidirectional-LSTM Model},
author = {Hang Thi-Thuy Do and Huy Duc Huynh and Kiet Van Nguyen and Ngan Luu-Thuy Nguyen and Anh Gia-Tuan Nguyen},
journal= {arXiv preprint arXiv:1911.03648},
year = {2019}
}