中文
相关论文

相关论文: KOLD: Korean Offensive Language Dataset

200 篇论文

Detecting offensive language on Twitter has many applications ranging from detecting/predicting bullying to measuring polarization. In this paper, we focus on building a large Arabic offensive tweet dataset. We introduce a method for…

计算与语言 · 计算机科学 2021-03-11 Hamdy Mubarak , Ammar Rashed , Kareem Darwish , Younes Samih , Ahmed Abdelali

The proliferation of hate speech has inflicted significant societal harm, with its intensity and directionality closely tied to specific targets and arguments. In recent years, numerous machine learning-based methods have been developed to…

计算与语言 · 计算机科学 2025-07-16 Zewen Bai , Liang Yang , Shengdi Yin , Yuanyuan Sun , Hongfei Lin

This paper introduces EVOKE (Emotion Vocabulary of Korean and English), a Korean-English parallel dataset of emotion words. The dataset offers comprehensive coverage of emotion words in each language, in addition to many-to-many…

计算与语言 · 计算机科学 2026-04-13 Yoonwon Jung , Hagyeong Shin , Benjamin K. Bergen

Scaling semantic parsing models for task-oriented dialog systems to new languages is often expensive and time-consuming due to the lack of available datasets. Available datasets suffer from several shortcomings: a) they contain few…

计算与语言 · 计算机科学 2021-01-28 Haoran Li , Abhinav Arora , Shuohui Chen , Anchit Gupta , Sonal Gupta , Yashar Mehdad

Instruction tuning has emerged as a powerful technique, significantly boosting zero-shot performance on unseen tasks. While recent work has explored cross-lingual generalization by applying instruction tuning to multilingual models,…

计算与语言 · 计算机科学 2024-06-14 Janghoon Han , Changho Lee , Joongbo Shin , Stanley Jungkyu Choi , Honglak Lee , Kynghoon Bae

Online hatred is a growing concern on many social media platforms. To address this issue, different social media platforms have introduced moderation policies for such content. They also employ moderators who can check the posts violating…

计算与语言 · 计算机科学 2021-12-01 Mithun Das , Somnath Banerjee , Punyajoy Saha

Online abusive behavior is an important issue that breaks the cohesiveness of online social communities and even raises public safety concerns in our societies. Motivated by this rising issue, researchers have proposed, collected, and…

社会与信息网络 · 计算机科学 2020-06-25 Md Rabiul Awal , Rui Cao , Roy Ka-Wei Lee , Sandra Mitrović

Word order variances generally exist in different languages. In this paper, we hypothesize that cross-lingual models that fit into the word order of the source language might fail to handle target languages. To verify this hypothesis, we…

计算与语言 · 计算机科学 2020-12-09 Zihan Liu , Genta Indra Winata , Samuel Cahyawijaya , Andrea Madotto , Zhaojiang Lin , Pascale Fung

Offensive language detection is one of the most challenging problem in the natural language processing field, being imposed by the rising presence of this phenomenon in online social media. This paper describes our Transformer-based…

计算与语言 · 计算机科学 2020-10-28 Mircea-Adrian Tanase , Dumitru-Clementin Cercel , Costin-Gabriel Chiru

The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely…

计算与语言 · 计算机科学 2025-02-24 Shiza Ali , Jeremy Blackburn , Gianluca Stringhini

With the exponential rise in user-generated web content on social media, the proliferation of abusive languages towards an individual or a group across the different sections of the internet is also rapidly increasing. It is very…

计算与语言 · 计算机科学 2021-03-24 Prashant Kapil , Asif Ekbal

The digital age has expanded social media and online forums, allowing free expression for nearly 45% of the global population. Yet, it has also fueled online harassment, bullying, and harmful behaviors like hate speech and toxic comments…

计算与语言 · 计算机科学 2026-03-12 Vuong M. Ngo , Cach N. Dang , Kien V. Nguyen , Mark Roantree

Social media sites such as YouTube and Facebook have become an integral part of everyone's life and in the last few years, hate speech in the social media comment section has increased rapidly. Detection of hate speech on social media…

计算与语言 · 计算机科学 2020-12-18 Nauros Romim , Mosahed Ahmed , Hriteshwar Talukder , Md Saiful Islam

As language models are often deployed as chatbot assistants, it becomes a virtue for models to engage in conversations in a user's first language. While these models are trained on a wide range of languages, a comprehensive evaluation of…

计算与语言 · 计算机科学 2024-06-18 Seongbo Jang , Seonghyeon Lee , Hwanjo Yu

Offensive content moderation is vital in social media platforms to support healthy online discussions. However, their prevalence in codemixed Dravidian languages is limited to classifying whole comments without identifying part of it…

When tasked with supporting multiple languages for a given problem, two approaches have arisen: training a model for each language with the annotation budget divided equally among them, and training on a high-resource language followed by…

计算与语言 · 计算机科学 2022-04-05 Joel Ruben Antony Moniz , Barun Patra , Matthew R. Gormley

Creating multilingual task-oriented dialogue (TOD) agents is challenging due to the high cost of training data acquisition. Following the research trend of improving training data efficiency, we show for the first time, that in-context…

A lack of demographic context in existing toxic speech datasets limits our understanding of how different age groups communicate online. In collaboration with funk, a German public service content network, this research introduces the first…

计算与语言 · 计算机科学 2025-09-01 Jan Fillies , Michael Peter Hoffmann , Rebecca Reichel , Roman Salzwedel , Sven Bodemer , Adrian Paschke

Social media platforms have recently seen an increase in the occurrence of hate speech discourse which has led to calls for improved detection methods. Most of these rely on annotated data, keywords, and a classification technique. While…

计算与语言 · 计算机科学 2017-11-29 Jherez Taylor , Melvyn Peignon , Yi-Shin Chen

Over the past decade, we have seen exponential growth in online content fueled by social media platforms. Data generation of this scale comes with the caveat of insurmountable offensive content in it. The complexity of identifying offensive…