中文

2023 年 ICON 印度语系性别虐待检测共享任务概述

计算与语言 2024-01-09 v1 机器学习

摘要

本文报告了 2023 年 ICON 关于印度语系性别虐待检测的发现。该共享任务旨在检测在线文本中的性别虐待。该共享任务是作为 2023 年 ICON 的一部分进行的,基于印地语、泰米尔语和印度英语方言的新数据集。参与者被分配了三个子任务,训练数据集包含来自 Twitter 的约 6500 条帖子。测试集提供了约 1200 条帖子。该共享任务共收到 9 份注册。各子任务的最佳 F-1 分数分别为:子任务 1 为 0.616,子任务 2 为 0.572,子任务 3 为 0.616 和 0.582。由于其主题原因,本文包含仇恨内容的示例。

关键词

引用

@article{arxiv.2401.03677,
  title  = {Overview of the 2023 ICON Shared Task on Gendered Abuse Detection in Indic Languages},
  author = {Aatman Vaidya and Arnav Arora and Aditya Joshi and Tarunima Prabhakar},
  journal= {arXiv preprint arXiv:2401.03677},
  year   = {2024}
}

备注

This paper has been accepted at 20th International Conference on Natural Language Processing (ICON), it is of 5 pages