中文
相关论文

相关论文: Directions in Abusive Language Training Data: Garb…

200 篇论文

In-game toxic language becomes the hot potato in the gaming industry and community. There have been several online game toxicity analysis frameworks and models proposed. However, it is still challenging to detect toxicity due to the nature…

计算与语言 · 计算机科学 2025-04-02 Yuanzhe Jia , Weixuan Wu , Feiqi Cao , Soyeon Caren Han

The issue of hate speech extends beyond the confines of the online realm. It is a problem with real-life repercussions, prompting most nations to formulate legal frameworks that classify hate speech as a punishable offence. These legal…

计算与语言 · 计算机科学 2024-12-10 Katerina Korre , John Pavlopoulos , Paolo Gajo , Alberto Barrón-Cedeño

Whereas much of the success of the current generation of neural language models has been driven by increasingly large training corpora, relatively little research has been dedicated to analyzing these massive sources of textual data. In…

计算与语言 · 计算机科学 2021-06-02 Alexandra Sasha Luccioni , Joseph D. Viviano

The problems of online hate speech and cyberbullying have significantly worsened since the increase in popularity of social media platforms such as YouTube and Twitter (X). Natural Language Processing (NLP) techniques have proven to provide…

计算与语言 · 计算机科学 2024-03-18 Sargam Yadav , Abhishek Kaushik , Kevin McDaid

As open-ended human-chatbot interaction becomes commonplace, sensitive content detection gains importance. In this work, we propose a two stage semi-supervised approach to bootstrap large-scale data for automatic sensitive language…

计算与语言 · 计算机科学 2018-12-03 Chandra Khatri , Behnam Hedayatnia , Rahul Goel , Anushree Venkatesh , Raefer Gabriel , Arindam Mandal

The ever growing usage of social media in the recent years has had a direct impact on the increased presence of hate speech and offensive speech in online platforms. Research on effective detection of such content has mainly focused on…

计算与语言 · 计算机科学 2022-05-11 Erida Nurce , Jorgel Keci , Leon Derczynski

Linguistics has been instrumental in developing a deeper understanding of human nature. Words are indispensable to bequeath the thoughts, emotions, and purpose of any human interaction, and critically analyzing these words can elucidate the…

计算与语言 · 计算机科学 2021-07-22 Tushar Sarkar , Nishant Rajadhyaksha

Hate speech has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. Multiple approaches have been developed to detect hate speech using artificial intelligence, but a generalized model is…

计算与语言 · 计算机科学 2024-10-10 Gautam Kishore Shahi , Tim A. Majchrzak

Online platforms have become an increasingly prominent means of communication. Despite the obvious benefits to the expanded distribution of content, the last decade has resulted in disturbing toxic communication, such as cyberbullying and…

社会与信息网络 · 计算机科学 2023-09-04 Amit Sheth , Valerie L. Shalin , Ugur Kursuncu

Data contamination has garnered increased attention in the era of large language models (LLMs) due to the reliance on extensive internet-derived training corpora. The issue of training corpus overlap with evaluation benchmarks--referred to…

计算与语言 · 计算机科学 2024-06-24 Chunyuan Deng , Yilun Zhao , Yuzhao Heng , Yitong Li , Jiannan Cao , Xiangru Tang , Arman Cohan

The spread of hate speech on social media space is currently a serious issue. The undemanding access to the enormous amount of information being generated on these platforms has led people to post and react with toxic content that…

计算与语言 · 计算机科学 2022-09-13 Abhishek Velankar , Hrushikesh Patil , Raviraj Joshi

As offensive language has become a rising issue for online communities and social media platforms, researchers have been investigating ways of coping with abusive content and developing systems to detect its different types: cyberbullying,…

计算与语言 · 计算机科学 2020-03-19 Zeses Pitenis , Marcos Zampieri , Tharindu Ranasinghe

Supervised machine learning, in which models are automatically derived from labeled training data, is only as good as the quality of that data. This study builds on prior work that investigated to what extent 'best practices' around…

机器学习 · 计算机科学 2021-07-07 R. Stuart Geiger , Dominique Cope , Jamie Ip , Marsha Lotosh , Aayush Shah , Jenny Weng , Rebekah Tang

Cyberbullying is a problem in today's ubiquitous online communities. Filtering it out of online conversations has proven a challenge, and efforts have led to the creation of many different datasets, all offered as resources to train…

计算与语言 · 计算机科学 2020-09-03 Khoury Richard , Larochelle Marc-André

Implicit discourse relation classification is one of the most challenging and important tasks in discourse parsing, due to the lack of connective as strong linguistic cues. A principle bottleneck to further improvement is the shortage of…

计算与语言 · 计算机科学 2019-04-16 Wei Shi , Frances Yung , Vera Demberg

The increased proliferation of abusive content on social media platforms has a negative impact on online users. The dread, dislike, discomfort, or mistrust of lesbian, gay, transgender or bisexual persons is defined as…

The goal of hate speech detection is to filter negative online content aiming at certain groups of people. Due to the easy accessibility of social media platforms it is crucial to protect everyone which requires building hate speech…

计算与语言 · 计算机科学 2022-01-19 Irina Bigoulaeva , Viktor Hangya , Iryna Gurevych , Alexander Fraser

The success of AI models relies on the availability of large, diverse, and high-quality datasets, which can be challenging to obtain due to data scarcity, privacy concerns, and high costs. Synthetic data has emerged as a promising solution…

计算与语言 · 计算机科学 2024-08-13 Ruibo Liu , Jerry Wei , Fangyu Liu , Chenglei Si , Yanzhe Zhang , Jinmeng Rao , Steven Zheng , Daiyi Peng , Diyi Yang , Denny Zhou , Andrew M. Dai

Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio…