English
Related papers

Related papers: Multilingual Twitter Corpus and Baselines for Eval…

200 papers

Most existing work on adversarial data generation focuses on English. For example, PAWS (Paraphrase Adversaries from Word Scrambling) consists of challenging English paraphrase identification pairs from Wikipedia and Quora. We remedy this…

Computation and Language · Computer Science 2019-09-02 Yinfei Yang , Yuan Zhang , Chris Tar , Jason Baldridge

Pretrained multilingual models exhibit the same social bias as models processing English texts. This systematic review analyzes emerging research that extends bias evaluation and mitigation approaches into multilingual and non-English…

Computation and Language · Computer Science 2025-09-08 Lance Calvin Lim Gamboa , Yue Feng , Mark Lee

Digital dehumanization, although a critical issue, remains largely overlooked within the field of computational linguistics and Natural Language Processing. The prevailing approach in current research concentrating primarily on a single…

Computation and Language · Computer Science 2025-10-22 Dennis Assenmacher , Paloma Piot , Katarina Laken , David Jurgens , Claudia Wagner

We construct Global Voices, a multilingual dataset for evaluating cross-lingual summarization methods. We extract social-network descriptions of Global Voices news articles to cheaply collect evaluation data for into-English and…

Computation and Language · Computer Science 2020-06-16 Khanh Nguyen , Hal Daumé

Bias mitigation approaches reduce models' dependence on sensitive features of data, such as social group tokens (SGTs), resulting in equal predictions across the sensitive features. In hate speech detection, however, equalizing model…

Computation and Language · Computer Science 2021-08-05 Aida Mostafazadeh Davani , Ali Omrani , Brendan Kennedy , Mohammad Atari , Xiang Ren , Morteza Dehghani

The damaging effects of hate speech on social media are evident during the last few years, and several organizations, researchers and social media platforms tried to harness them in various ways. Despite these efforts, social media users…

Information Retrieval · Computer Science 2020-05-04 Polychronis Charitidis , Stavros Doropoulos , Stavros Vologiannidis , Ioannis Papastergiou , Sophia Karakeva

As a result of social network popularity, in recent years, hate speech phenomenon has significantly increased. Due to its harmful effect on minority groups as well as on large communities, there is a pressing need for hate speech detection…

Computation and Language · Computer Science 2019-12-13 Kristian Miok , Dong Nguyen-Doan , Blaž Škrlj , Daniela Zaharie , Marko Robnik-Šikonja

With the ever-increasing cases of hate spread on social media platforms, it is critical to design abuse detection mechanisms to proactively avoid and control such incidents. While there exist methods for hate speech detection, they…

Computation and Language · Computer Science 2020-01-17 Pinkesh Badjatiya , Manish Gupta , Vasudeva Varma

Offensive content is pervasive in social media and a reason for concern to companies and government organizations. Several studies have been recently published investigating methods to detect the various forms of such content (e.g. hate…

Computation and Language · Computer Science 2021-05-21 Tharindu Ranasinghe , Marcos Zampieri

While the impact of social biases in language models has been recognized, prior methods for bias evaluation have been limited to binary association tests on small datasets, limiting our understanding of bias complexities. This paper…

Computation and Language · Computer Science 2025-05-27 Marta Marchiori Manerba , Karolina Stańczak , Riccardo Guidotti , Isabelle Augenstein

Algorithmic hate speech detection faces significant challenges due to the diverse definitions and datasets used in research and practice. Social media platforms, legal frameworks, and institutions each apply distinct yet overlapping…

Computation and Language · Computer Science 2025-03-10 Jan Fillies , Adrian Paschke

Well-annotated data is a prerequisite for good Natural Language Processing models. Too often, though, annotation decisions are governed by optimizing time or annotator agreement. We make a case for nuanced efforts in an interdisciplinary…

Computation and Language · Computer Science 2022-10-31 Federico Bianchi , Stefanie Anja Hills , Patricia Rossini , Dirk Hovy , Rebekah Tromble , Nava Tintarev

Human biases are ubiquitous but not uniform: disparities exist across linguistic, cultural, and societal borders. As large amounts of recent literature suggest, language models (LMs) trained on human data can reflect and often amplify the…

Computation and Language · Computer Science 2023-10-27 Anjishnu Mukherjee , Chahat Raj , Ziwei Zhu , Antonios Anastasopoulos

Suicidal ideation is a serious health problem affecting millions of people worldwide. Social networks provide information about these mental health problems through users' emotional expressions. We propose a multilingual model leveraging…

Computation and Language · Computer Science 2024-12-23 Rodolfo Zevallos , Annika Schoene , John E. Ortega

Automatic identification of hateful and abusive content is vital in combating the spread of harmful online content and its damaging effects. Most existing works evaluate models by examining the generalization error on train-test splits on…

Computation and Language · Computer Science 2025-04-07 Lanqin Yuan , Marian-Andrei Rizoiu

In this work we target the problem of hate speech detection in multimodal publications formed by a text and an image. We gather and annotate a large scale dataset from Twitter, MMHS150K, and propose different models that jointly analyze…

Computer Vision and Pattern Recognition · Computer Science 2019-10-10 Raul Gomez , Jaume Gibert , Lluis Gomez , Dimosthenis Karatzas

Although there is an unprecedented effort to provide adequate responses in terms of laws and policies to hate content on social media platforms, dealing with hatred online is still a tough problem. Tackling hate speech in the standard way…

Computation and Language · Computer Science 2019-10-09 Y. L. Chung , E. Kuzmenko , S. S. Tekiroglu , M. Guerini

Abusive language detection has become an increasingly important task as a means to tackle this type of harmful content in social media. There has been a substantial body of research developing models for determining if a social media post…

Computation and Language · Computer Science 2025-08-19 Raneem Alharthi , Rajwa Alharthi , Aiqi Jiang , Arkaitz Zubiaga

In this paper, we investigate the issue of hate speech by presenting a novel task of translating hate speech into non-hate speech text while preserving its meaning. As a case study, we use Spanish texts. We provide a dataset and several…

Computation and Language · Computer Science 2023-06-05 Yevhen Kostiuk , Atnafu Lambebo Tonja , Grigori Sidorov , Olga Kolesnikova

Cyberbullying is a pervasive problem in online communities. To identify cyberbullying cases in large-scale social networks, content moderators depend on machine learning classifiers for automatic cyberbullying detection. However, existing…

Social and Information Networks · Computer Science 2020-04-07 Caleb Ziems , Ymir Vigfusson , Fred Morstatter