English
Related papers

Related papers: KOLD: Korean Offensive Language Dataset

200 papers

The number of increased social media users has led to a lot of people misusing these platforms to spread offensive content and use hate speech. Manual tracking the vast amount of posts is impractical so it is necessary to devise automated…

Computation and Language · Computer Science 2022-02-08 Arka Mitra , Priyanshu Sankhala

Health departments have been deploying text classification systems for the early detection of foodborne illness complaints in social media documents such as Yelp restaurant reviews. Current systems have been successfully applied for…

Computation and Language · Computer Science 2020-10-13 Ziyi Liu , Giannis Karamanolakis , Daniel Hsu , Luis Gravano

Public figures receive a disproportionate amount of abuse on social media, impacting their active participation in public life. Automated systems can identify abuse at scale but labelling training data is expensive, complex and potentially…

In this study, we introduce KOPL, a novel framework for handling Korean OOV words with Phoneme representation Learning. Our work is based on the linguistic property of Korean as a phonemic script, the high correlation between phonemes and…

Computation and Language · Computer Science 2025-07-08 Nayeon Kim , Eojin Jeon , Jun-Hyung Park , SangKeun Lee

Although offensive language continually evolves over time, even recent studies using LLMs have predominantly relied on outdated datasets and rarely evaluated the generalization ability on unseen texts. In this study, we constructed a…

Computation and Language · Computer Science 2025-09-19 Seunguk Yu , Jungmin Yun , Jinhee Jang , Youngbin Kim

Existing rhetorical understanding and generation datasets or corpora primarily focus on single coarse-grained categories or fine-grained categories, neglecting the common interrelations between different rhetorical devices by treating them…

Computation and Language · Computer Science 2024-10-01 Nuowei Liu , Xinhao Chen , Hongyi Wu , Changzhi Sun , Man Lan , Yuanbin Wu , Xiaopeng Bai , Shaoguang Mao , Yan Xia

Language detoxification involves removing toxicity from offensive language. While a neutral-toxic paired dataset provides a straightforward approach for training detoxification models, creating such datasets presents several challenges: i)…

Computation and Language · Computer Science 2025-06-17 Minkyeong Jeon , Hyemin Jeong , Yerang Kim , Jiyoung Kim , Jae Hyeon Cho , Byung-Jun Lee

With the freedom of communication provided in online social media, hate speech has increasingly generated. This leads to cyber conflicts affecting social life at the individual and national levels. As a result, hateful content…

Computation and Language · Computer Science 2022-09-16 Khouloud Mnassri , Praboda Rajapaksha , Reza Farahbakhsh , Noel Crespi

In the digital era, effective identification and analysis of verbal attacks are essential for maintaining online civility and ensuring social security. However, existing research is limited by insufficient modeling of conversational…

Computation and Language · Computer Science 2026-01-13 Quan Zheng , Yuanhe Tian , Ming Wang , Yan Song

The ability to accurately detect and filter offensive content automatically is important to ensure a rich and diverse digital discourse. Trolling is a type of hurtful or offensive content that is prevalent in social media, but is…

Computers and Society · Computer Science 2020-08-04 Hitkul , Karmanya Aggarwal , Pakhi Bamdev , Debanjan Mahata , Rajiv Ratn Shah , Ponnurangam Kumaraguru

The dissemination of online hate speech can have serious negative consequences for individuals, online communities, and entire societies. This and the large volume of hateful online content prompted both practitioners', i.e., in content…

Computation and Language · Computer Science 2025-04-14 Julian Bäumler , Louis Blöcher , Lars-Joel Frey , Xian Chen , Markus Bayer , Christian Reuter

With the growing presence of multilingual users on social media, detecting abusive language in code-mixed text has become increasingly challenging. Code-mixed communication, where users seamlessly switch between English and their native…

Computation and Language · Computer Science 2025-05-01 Manish Pandey , Nageshwar Prasad Yadav , Mokshada Adduru , Sawan Rai

The paper presents the submission of the team indicnlp@kgp to the EACL 2021 shared task "Offensive Language Identification in Dravidian Languages." The task aimed to classify different offensive content types in 3 code-mixed Dravidian…

Computation and Language · Computer Science 2021-02-16 Kushal Kedia , Abhilash Nandy

In recent years, there has been a lot of focus on offensive content. The amount of offensive content generated by social media is increasing at an alarming rate. This created a greater need to address this issue than ever before. To address…

Computation and Language · Computer Science 2022-04-22 Shankar Biradar , Sunil Saumya

Hate speech detection models are only as good as the data they are trained on. Datasets sourced from social media suffer from systematic gaps and biases, leading to unreliable models with simplistic decision boundaries. Adversarial…

Computation and Language · Computer Science 2024-03-29 Janis Goldzycher , Paul Röttger , Gerold Schneider

Hate speech has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. Multiple approaches have been developed to detect hate speech using artificial intelligence, but a generalized model is…

Computation and Language · Computer Science 2024-10-10 Gautam Kishore Shahi , Tim A. Majchrzak

Detecting offensive language on social media is an important task. The ICWSM-2020 Data Challenge Task 2 is aimed at identifying offensive content using a crowd-sourced dataset containing 100k labelled tweets. The dataset, however, suffers…

Computation and Language · Computer Science 2020-12-08 Ruibo Liu , Guangxuan Xu , Soroush Vosoughi

Transformer-based models such as BERT, XLNET, and XLM-R have achieved state-of-the-art performance across various NLP tasks including the identification of offensive language and hate speech, an important problem in social media. In this…

Computation and Language · Computer Science 2021-09-14 Diptanu Sarkar , Marcos Zampieri , Tharindu Ranasinghe , Alexander Ororbia

Sentiment analysis is the most basic NLP task to determine the polarity of text data. There has been a significant amount of work in the area of multilingual text as well. Still hate and offensive speech detection faces a challenge due to…

Computation and Language · Computer Science 2021-11-02 Abhishek Velankar , Hrushikesh Patil , Amol Gore , Shubham Salunke , Raviraj Joshi

We present the results and the main findings of SemEval-2019 Task 6 on Identifying and Categorizing Offensive Language in Social Media (OffensEval). The task was based on a new dataset, the Offensive Language Identification Dataset (OLID),…

Computation and Language · Computer Science 2019-04-30 Marcos Zampieri , Shervin Malmasi , Preslav Nakov , Sara Rosenthal , Noura Farra , Ritesh Kumar
‹ Prev 1 3 4 5 6 7 10 Next ›