中文
相关论文

相关论文: Necessity and Sufficiency for Explaining Text Clas…

200 篇论文

Large language models (LLMs) excel in many diverse applications beyond language generation, e.g., translation, summarization, and sentiment analysis. One intriguing application is in text classification. This becomes pertinent in the realm…

计算与语言 · 计算机科学 2024-03-14 Tharindu Kumarage , Amrita Bhattacharjee , Joshua Garland

Discriminatory language and biases are often present in hate speech during conversations, which usually lead to negative impacts on targeted groups such as those based on race, gender, and religion. To tackle this issue, we propose an…

计算与语言 · 计算机科学 2023-07-21 Shaina Raza , Chen Ding , Deval Pandya

Today, the internet is an integral part of our daily lives, enabling people to be more connected than ever before. However, this greater connectivity and access to information increase exposure to harmful content such as cyber-bullying and…

社会与信息网络 · 计算机科学 2023-10-31 Lanqin Yuan , Tianyu Wang , Gabriela Ferraro , Hanna Suominen , Marian-Andrei Rizoiu

Feature attribution methods, proposed recently, help users interpret the predictions of complex models. Our approach integrates feature attributions into the objective function to allow machine learning practitioners to incorporate priors…

计算与语言 · 计算机科学 2019-06-21 Frederick Liu , Besim Avci

Warning: This paper contains examples of the language that some people may find offensive. Detecting and reducing hateful, abusive, offensive comments is a critical and challenging task on social media. Moreover, few studies aim to mitigate…

计算与语言 · 计算机科学 2023-12-21 Neeraj Kumar Singh , Koyel Ghosh , Joy Mahapatra , Utpal Garain , Apurbalal Senapati

Hate speech has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. Multiple approaches have been developed to detect hate speech using artificial intelligence, but a generalized model is…

计算与语言 · 计算机科学 2024-10-10 Gautam Kishore Shahi , Tim A. Majchrzak

Some users of social media are spreading racist, sexist, and otherwise hateful content. For the purpose of training a hate speech detection system, the reliability of the annotations is crucial, but there is no universally agreed-upon…

计算与语言 · 计算机科学 2017-01-30 Björn Ross , Michael Rist , Guillermo Carbonell , Benjamin Cabrera , Nils Kurowsky , Michael Wojatzki

Given a machine learning (ML) model and a prediction, explanations can be defined as sets of features which are sufficient for the prediction. In some applications, and besides asking for an explanation, it is also critical to understand…

机器学习 · 计算机科学 2023-02-08 Xuanxiang Huang , Martin C. Cooper , Antonio Morgado , Jordi Planes , Joao Marques-Silva

Feature attribution a.k.a. input salience methods which assign an importance score to a feature are abundant but may produce surprisingly different results for the same model on the same input. While differences are expected if disparate…

计算与语言 · 计算机科学 2022-11-10 Jasmijn Bastings , Sebastian Ebert , Polina Zablotskaia , Anders Sandholm , Katja Filippova

Readability assessment aims to automatically classify text by the level appropriate for learning readers. Traditional approaches to this task utilize a variety of linguistically motivated features paired with simple machine learning models.…

计算与语言 · 计算机科学 2020-08-04 Tovly Deutsch , Masoud Jasbi , Stuart Shieber

In a hate speech detection model, we should consider two critical aspects in addition to detection performance-bias and explainability. Hate speech cannot be identified based solely on the presence of specific words: the model should be…

计算与语言 · 计算机科学 2022-11-02 Jiyun Kim , Byounghan Lee , Kyung-Ah Sohn

This work proposes a contextualised detection framework for implicitly hateful speech, implemented as a multi-agent system comprising a central Moderator Agent and dynamically constructed Community Agents representing specific demographic…

计算与语言 · 计算机科学 2026-01-28 Ewelina Gajewska , Katarzyna Budzynska , Jarosław A Chudziak

Feature importance scores are ubiquitous tools for understanding the predictions of machine learning models. However, many popular attribution methods suffer from high instability due to random sampling. Leveraging novel ideas from…

机器学习 · 统计学 2025-07-08 Jeremy Goldwasser , Giles Hooker

The rise of social networks has not only facilitated communication but also allowed the spread of harmful content. Although significant advances have been made in detecting toxic language in textual data, the exploration of concept-based…

计算与语言 · 计算机科学 2025-12-16 Samarth Garg , Divya Singh , Deeksha Varshney , Mamta

Hate speech frequently appears on social media platforms and urgently needs to be effectively controlled. Alleviating the bias caused by hate speech can help resolve various ethical issues. Although existing research has constructed several…

计算与语言 · 计算机科学 2025-08-27 Hongyan Wu , Zhengming Chen , Zijian Li , Nankai Lin , Lianxi Wang , Shengyi Jiang , Aimin Yang

Progress on many Natural Language Processing (NLP) tasks, such as text classification, is driven by objective, reproducible and scalable evaluation via publicly available benchmarks. However, these are not always representative of…

计算与语言 · 计算机科学 2022-11-11 Viktor Schlegel , Erick Mendez-Guzman , Riza Batista-Navarro

The detection of hate speech online has become an important task, as offensive language such as hurtful, obscene and insulting content can harm marginalized people or groups. This paper presents TU Berlin team experiments and results on the…

计算与语言 · 计算机科学 2022-01-13 Salar Mohtaj , Vera Schmitt , Sebastian Möller

Hate speech detection is a common downstream application of natural language processing (NLP) in the real world. In spite of the increasing accuracy, current data-driven approaches could easily learn biases from the imbalanced data…

计算与语言 · 计算机科学 2022-09-22 Yi Cai , Arthur Zimek , Gerhard Wunder , Eirini Ntoutsi

Hate speech is one of the main threats posed by the widespread use of social networks, despite efforts to limit it. Although attention has been devoted to this issue, the lack of datasets and case studies centered around scarcely…

计算与语言 · 计算机科学 2024-10-11 Camilla Casula , Sara Tonelli

With the spreading of hate speech on social media in recent years, automatic detection of hate speech is becoming a crucial task and has attracted attention from various communities. This task aims to recognize online posts (e.g., tweets)…

计算与语言 · 计算机科学 2022-04-15 Jiaxuan Li , Yue Ning