English
Related papers

Related papers: Can Interpretability Layouts Influence Human Perce…

200 papers

Large language models (LLMs) are increasingly used in decision-making tasks like r\'esum\'e screening and content moderation, giving them the power to amplify or suppress certain perspectives. While previous research has identified…

Computation and Language · Computer Science 2025-05-28 Naba Rizvi , Harper Strickland , Saleha Ahmedi , Aekta Kallepalli , Isha Khirwadkar , William Wu , Imani N. S. Munyaka , Nedjma Ousidhoum

Machine learning (ML) interpretability techniques can reveal undesirable patterns in data that models exploit to make predictions--potentially causing harms once deployed. However, how to take action to address these patterns is not always…

Native speakers can judge whether a sentence is an acceptable instance of their language. Acceptability provides a means of evaluating whether computational language models are processing language in a human-like manner. We test the ability…

Computation and Language · Computer Science 2019-10-11 Wang Jing , M. A. Kelly , David Reitter

Social media platforms are deploying machine learning based offensive language classification systems to combat hateful, racist, and other forms of offensive speech at scale. However, despite their real-world deployment, we do not yet…

Computation and Language · Computer Science 2022-03-23 Jonathan Rusert , Zubair Shafiq , Padmini Srinivasan

Automated counter-narratives (CN) offer a promising strategy for mitigating online hate speech, yet concerns about their affective tone, accessibility, and ethical risks remain. We propose a framework for evaluating Large Language Model…

Computation and Language · Computer Science 2025-06-05 Mikel K. Ngueajio , Flor Miriam Plaza-del-Arco , Yi-Ling Chung , Danda B. Rawat , Amanda Cercas Curry

The increasing adoption of machine learning tools has led to calls for accountability via model interpretability. But what does it mean for a machine learning model to be interpretable by humans, and how can this be assessed? We focus on…

Machine Learning · Computer Science 2019-08-06 Dylan Slack , Sorelle A. Friedler , Carlos Scheidegger , Chitradeep Dutta Roy

Social media platforms provide users the freedom of expression and a medium to exchange information and express diverse opinions. Unfortunately, this has also resulted in the growth of abusive content with the purpose of discriminating…

Computation and Language · Computer Science 2021-07-01 Sohail Akhtar , Valerio Basile , Viviana Patti

Hate speech detection is a socially sensitive and inherently subjective task, with judgments often varying based on personal traits. While prior work has examined how socio-demographic factors influence annotation, the impact of personality…

Computation and Language · Computer Science 2025-06-11 Shuzhou Yuan , Ercong Nie , Mario Tawfelis , Helmut Schmid , Hinrich Schütze , Michael Färber

Studies across many disciplines have shown that lexical choice can affect audience perception. For example, how users describe themselves in a social media profile can affect their perceived socio-economic status. However, we lack general…

Machine Learning · Computer Science 2018-11-16 Zhao Wang , Aron Culotta

The Rashomon effect describes the observation that in machine learning (ML) multiple models often achieve similar predictive performance while explaining the underlying relationships in different ways. This observation holds even for…

Machine Learning · Computer Science 2025-05-13 Julian Rosenberger , Philipp Schröppel , Sven Kruschel , Mathias Kraus , Patrick Zschech , Maximilian Förster

Detecting harmful content is a crucial task in the landscape of NLP applications for Social Good, with hate speech being one of its most dangerous forms. But what do we mean by hate speech, how can we define it, and how does prompting…

Computation and Language · Computer Science 2025-06-24 Matteo Melis , Gabriella Lapesa , Dennis Assenmacher

In this paper we investigate the explainability of transformer models and their plausibility for hate speech and counter speech detection. We compare representatives of four different explainability approaches, i.e., gradient-based,…

Machine Learning · Computer Science 2024-07-31 Adrian Jaques Böck , Djordje Slijepčević , Matthias Zeppelzauer

Machine Learning (ML) is increasingly applied in real-life scenarios, raising concerns about bias in automatic decision making. We focus on bias as a notion of opinion exclusion, that stems from the direct application of traditional ML…

Machine Learning · Computer Science 2019-11-07 Agathe Balayn , Alessandro Bozzon

Amidst the rapid expansion of Machine Learning (ML) and Large Language Models (LLMs), understanding the semantics within their mechanisms is vital. Causal analyses define semantics, while gradient-based methods are essential to eXplainable…

Artificial Intelligence · Computer Science 2024-03-26 Yosuke Miyanishi , Minh Le Nguyen

Language models (LMs) are known to represent the perspectives of some social groups better than others, which may impact their performance, especially on subjective tasks such as content moderation and hate speech detection. To explore how…

Computation and Language · Computer Science 2024-06-19 Zihao He , Siyi Guo , Ashwin Rao , Kristina Lerman

Large Language Models (LLMs) have become essential for offensive language detection, yet their ability to handle annotation disagreement remains underexplored. Disagreement samples, which arise from subjective interpretations, pose a unique…

Computation and Language · Computer Science 2025-05-20 Junyu Lu , Kai Ma , Kaichun Wang , Kelaiti Xiao , Roy Ka-Wei Lee , Bo Xu , Liang Yang , Hongfei Lin

Recent efforts in Machine Learning (ML) interpretability have focused on creating methods for explaining black-box ML models. However, these methods rely on the assumption that simple approximations, such as linear models or decision-trees,…

Machine Learning · Computer Science 2019-06-13 Owen Lahav , Nicholas Mastronarde , Mihaela van der Schaar

While the interpretability of machine learning models is often equated with their mere syntactic comprehensibility, we think that interpretability goes beyond that, and that human interpretability should also be investigated from the point…

Machine Learning · Statistics 2021-09-14 Tomáš Kliegr , Štěpán Bahník , Johannes Fürnkranz

The fairness and trustworthiness of Large Language Models (LLMs) are receiving increasing attention. Implicit hate speech, which employs indirect language to convey hateful intentions, occupies a significant portion of practice. However,…

Computation and Language · Computer Science 2024-07-24 Min Zhang , Jianfeng He , Taoran Ji , Chang-Tien Lu

The dissemination of online hate speech can have serious negative consequences for individuals, online communities, and entire societies. This and the large volume of hateful online content prompted both practitioners', i.e., in content…

Computation and Language · Computer Science 2025-04-14 Julian Bäumler , Louis Blöcher , Lars-Joel Frey , Xian Chen , Markus Bayer , Christian Reuter