English
Related papers

Related papers: Examining Temporal Bias in Abusive Language Detect…

200 papers

Recent research has revealed that neural language models at scale suffer from poor temporal generalization capability, i.e., the language model pre-trained on static data from past years performs worse over time on emerging data. Existing…

Computation and Language · Computer Science 2022-11-01 Zhaochen Su , Zecheng Tang , Xinyan Guan , Juntao Li , Lijun Wu , Min Zhang

The spread of online hate has become a significant problem for newspapers that host comment sections. As a result, there is growing interest in using machine learning and natural language processing for (semi-) automated abusive language…

Computation and Language · Computer Science 2022-07-11 Lennart Justen , Kilian Müller , Marco Niemann , Jörg Becker

While social media offers freedom of self-expression, abusive language carry significant negative social impact. Driven by the importance of the issue, research in the automated detection of abusive language has witnessed growth and…

Computation and Language · Computer Science 2022-05-04 Wenjie Yin , Arkaitz Zubiaga

The problem of online threats and abuse could potentially be mitigated with a computational approach, where sources of abuse are better understood or identified through author profiling. However, abusive language constitutes a specific…

Computation and Language · Computer Science 2020-09-04 Isabelle van der Vegt , Bennett Kleinberg , Paul Gill

Textual data used to train large language models (LLMs) exhibits multifaceted bias manifestations encompassing harmful language and skewed demographic distributions. Regulations such as the European AI Act require identifying and mitigating…

It is a well-known fact that current AI-based language technology -- language models, machine translation systems, multilingual dictionaries and corpora -- focuses on the world's 2-3% most widely spoken languages. Recent research efforts…

Computation and Language · Computer Science 2023-07-26 Gábor Bella , Paula Helm , Gertraud Koch , Fausto Giunchiglia

Social bias in language - towards genders, ethnicities, ages, and other social groups - poses a problem with ethical impact for many NLP applications. Recent research has shown that machine learning models trained on respective data may not…

Computation and Language · Computer Science 2020-11-25 Maximilian Spliethöver , Henning Wachsmuth

It is well known that textual data on the internet and other digital platforms contain significant levels of bias and stereotypes. Although many such texts contain stereotypes and biases that inherently exist in natural language for reasons…

Computation and Language · Computer Science 2022-01-24 Ewoenam Kwaku Tokpo , Toon Calders

As the body of research on abusive language detection and analysis grows, there is a need for critical consideration of the relationships between different subtasks that have been grouped under this label. Based on work on hate speech,…

Computation and Language · Computer Science 2017-05-31 Zeerak Waseem , Thomas Davidson , Dana Warmsley , Ingmar Weber

Technology for language generation has advanced rapidly, spurred by advancements in pre-training large models on massive amounts of data and the need for intelligent agents to communicate in a natural manner. While techniques can…

Computation and Language · Computer Science 2021-06-24 Emily Sheng , Kai-Wei Chang , Premkumar Natarajan , Nanyun Peng

Recent studies in the field of Machine Translation (MT) and Natural Language Processing (NLP) have shown that existing models amplify biases observed in the training data. The amplification of biases in language technology has mainly been…

Computation and Language · Computer Science 2021-02-02 Eva Vanmassenhove , Dimitar Shterionov , Matthew Gwilliam

Translating from languages without productive grammatical gender like English into gender-marked languages is a well-known difficulty for machines. This difficulty is also due to the fact that the training data on which models are built…

Computation and Language · Computer Science 2020-06-11 Luisa Bentivogli , Beatrice Savoldi , Matteo Negri , Mattia Antonino Di Gangi , Roldano Cattoni , Marco Turchi

This study investigates the consequences of training language models on synthetic data generated by their predecessors, an increasingly prevalent practice given the prominence of powerful generative models. Diverging from the usual emphasis…

Computation and Language · Computer Science 2024-04-17 Yanzhu Guo , Guokan Shang , Michalis Vazirgiannis , Chloé Clavel

In this work, abusive language detection in online content is performed using Bidirectional Recurrent Neural Network (BiRNN) method. Here the main objective is to focus on various forms of abusive behaviors on Twitter and to detect whether…

Computation and Language · Computer Science 2020-10-01 Dincy Davis , Reena Murali , Remesh Babu

The remarkable progress in Natural Language Processing (NLP) brought about by deep learning, particularly with the recent advent of large pre-trained neural language models, is brought into scrutiny as several studies began to discuss and…

Computation and Language · Computer Science 2023-01-25 Anoop K. , Manjary P. Gangan , Deepak P. , Lajish V. L

Uses of pejorative expressions can be benign or actively empowering. When models for abuse detection misclassify these expressions as derogatory, they inadvertently censor productive conversations held by marginalized groups. One way to…

Computation and Language · Computer Science 2022-06-20 Jana Kurrek , Haji Mohammad Saleem , Derek Ruths

Biased associations have been a challenge in the development of classifiers for detecting toxic language, hindering both fairness and accuracy. As potential solutions, we investigate recently introduced debiasing methods for text…

Computation and Language · Computer Science 2021-02-02 Xuhui Zhou , Maarten Sap , Swabha Swayamdipta , Noah A. Smith , Yejin Choi

Temporal validity is an important property of text that is useful for many downstream applications, such as recommender systems, conversational AI, or story understanding. Existing benchmarking tasks often require models to identify the…

Computation and Language · Computer Science 2024-01-02 Georg Wenzel , Adam Jatowt

Hundreds of millions of people now interact with language models, with uses ranging from serving as a writing aid to informing hiring decisions. Yet these language models are known to perpetuate systematic racial prejudices, making their…

Computation and Language · Computer Science 2024-03-04 Valentin Hofmann , Pratyusha Ria Kalluri , Dan Jurafsky , Sharese King

Sociodemographic bias in language models (LMs) has the potential for harm when deployed in real-world settings. This paper presents a comprehensive survey of the past decade of research on sociodemographic bias in LMs, organized into a…

Computation and Language · Computer Science 2024-08-15 Vipul Gupta , Pranav Narayanan Venkit , Shomir Wilson , Rebecca J. Passonneau