2023 年 ICON 印度语系性别虐待检测共享任务概述
计算与语言
2024-01-09 v1 机器学习
摘要
本文报告了 2023 年 ICON 关于印度语系性别虐待检测的发现。该共享任务旨在检测在线文本中的性别虐待。该共享任务是作为 2023 年 ICON 的一部分进行的,基于印地语、泰米尔语和印度英语方言的新数据集。参与者被分配了三个子任务,训练数据集包含来自 Twitter 的约 6500 条帖子。测试集提供了约 1200 条帖子。该共享任务共收到 9 份注册。各子任务的最佳 F-1 分数分别为:子任务 1 为 0.616,子任务 2 为 0.572,子任务 3 为 0.616 和 0.582。由于其主题原因,本文包含仇恨内容的示例。
引用
@article{arxiv.2401.03677,
title = {Overview of the 2023 ICON Shared Task on Gendered Abuse Detection in Indic Languages},
author = {Aatman Vaidya and Arnav Arora and Aditya Joshi and Tarunima Prabhakar},
journal= {arXiv preprint arXiv:2401.03677},
year = {2024}
}
备注
This paper has been accepted at 20th International Conference on Natural Language Processing (ICON), it is of 5 pages