中文
相关论文

相关论文: Bridging Fairness and Explainability: Can Input-Ba…

200 篇论文

In order to build reliable and trustworthy NLP applications, models need to be both fair across different demographics and explainable. Usually these two objectives, fairness and explainability, are optimized and/or examined independently…

计算与语言 · 计算机科学 2023-11-14 Stephanie Brandl , Emanuele Bugliarello , Ilias Chalkidis

Hate speech detection is a common downstream application of natural language processing (NLP) in the real world. In spite of the increasing accuracy, current data-driven approaches could easily learn biases from the imbalanced data…

计算与语言 · 计算机科学 2022-09-22 Yi Cai , Arthur Zimek , Gerhard Wunder , Eirini Ntoutsi

Motivations for methods in explainable artificial intelligence (XAI) often include detecting, quantifying and mitigating bias, and contributing to making machine learning models fairer. However, exactly how an XAI method can help in…

计算与语言 · 计算机科学 2022-06-09 Esma Balkir , Svetlana Kiritchenko , Isar Nejadgholi , Kathleen C. Fraser

While machine learning models have achieved unprecedented success in real-world applications, they might make biased/unfair decisions for specific demographic groups and hence result in discriminative outcomes. Although research efforts…

机器学习 · 计算机科学 2022-12-08 Yuying Zhao , Yu Wang , Tyler Derr

Hate speech detection is a crucial area of research in natural language processing, essential for ensuring online community safety. However, detecting implicit hate speech, where harmful intent is conveyed in subtle or indirect ways,…

计算与语言 · 计算机科学 2025-04-17 Yumin Kim , Hwanhee Lee

In this paper we investigate the explainability of transformer models and their plausibility for hate speech and counter speech detection. We compare representatives of four different explainability approaches, i.e., gradient-based,…

机器学习 · 计算机科学 2024-07-31 Adrian Jaques Böck , Djordje Slijepčević , Matthias Zeppelzauer

The advent of social media has given rise to numerous ethical challenges, with hate speech among the most significant concerns. Researchers are attempting to tackle this problem by leveraging hate-speech detection and employing language…

计算与语言 · 计算机科学 2023-05-31 Pranath Reddy Kumbam , Sohaib Uddin Syed , Prashanth Thamminedi , Suhas Harish , Ian Perera , Bonnie J. Dorr

As NLP models become more integrated with the everyday lives of people, it becomes important to examine the social effect that the usage of these systems has. While these models understand language and have increased accuracy on difficult…

计算与语言 · 计算机科学 2022-04-21 Rajas Bansal

Language models are the new state-of-the-art natural language processing (NLP) models and they are being increasingly used in many NLP tasks. Even though there is evidence that language models are biased, the impact of that bias on the…

计算与语言 · 计算机科学 2024-04-29 Fatma Elsafoury , Stamos Katsigiannis

Machine learning models in safety-critical settings like healthcare are often blackboxes: they contain a large number of parameters which are not transparent to users. Post-hoc explainability methods where a simple, human-interpretable…

机器学习 · 计算机科学 2022-06-03 Aparna Balagopalan , Haoran Zhang , Kimia Hamidieh , Thomas Hartvigsen , Frank Rudzicz , Marzyeh Ghassemi

The fairness and trustworthiness of Large Language Models (LLMs) are receiving increasing attention. Implicit hate speech, which employs indirect language to convey hateful intentions, occupies a significant portion of practice. However,…

计算与语言 · 计算机科学 2024-07-24 Min Zhang , Jianfeng He , Taoran Ji , Chang-Tien Lu

In a hate speech detection model, we should consider two critical aspects in addition to detection performance-bias and explainability. Hate speech cannot be identified based solely on the presence of specific words: the model should be…

计算与语言 · 计算机科学 2022-11-02 Jiyun Kim , Byounghan Lee , Kyung-Ah Sohn

Human biases have been shown to influence the performance of models and algorithms in various fields, including Natural Language Processing. While the study of this phenomenon is garnering focus in recent years, the available resources are…

计算与语言 · 计算机科学 2024-08-15 Ana Sofia Evans , Helena Moniz , Luísa Coheur

Although social media platforms are a prominent arena for users to engage in interpersonal discussions and express opinions, the facade and anonymity offered by social media may allow users to spew hate speech and offensive content. Given…

计算与语言 · 计算机科学 2024-05-09 Ayushi Nirmal , Amrita Bhattacharjee , Paras Sheth , Huan Liu

Natural Language Processing (NLP) systems learn harmful societal biases that cause them to amplify inequality as they are deployed in more and more situations. To guide efforts at debiasing these systems, the NLP community relies on a…

计算与语言 · 计算机科学 2021-06-09 Seraphina Goldfarb-Tarrant , Rebecca Marchant , Ricardo Muñoz Sanchez , Mugdha Pandya , Adam Lopez

Debiasing methods in NLP models traditionally focus on isolating information related to a sensitive attribute (e.g., gender or race). We instead argue that a favorable debiasing method should use sensitive information 'fairly,' with…

计算与语言 · 计算机科学 2023-10-24 Bodhisattwa Prasad Majumder , Zexue He , Julian McAuley

Recent research at the intersection of AI explainability and fairness has focused on how explanations can improve human-plus-AI task performance as assessed by fairness measures. We propose to characterize what constitutes an explanation…

计算与语言 · 计算机科学 2023-10-24 Tin Nguyen , Jiannan Xu , Aayushi Roy , Hal Daumé , Marine Carpuat

There have been remarkable breakthroughs in Machine Learning and Artificial Intelligence, notably in the areas of Natural Language Processing and Deep Learning. Additionally, hate speech detection in dialogues has been gaining popularity…

计算与语言 · 计算机科学 2023-06-02 Durgesh Nandini , Ute Schmid

The automatic detection of hate speech online is an active research area in NLP. Most of the studies to date are based on social media datasets that contribute to the creation of hate speech detection models trained on them. However, data…

计算与语言 · 计算机科学 2023-07-06 Dimosthenis Antypas , Jose Camacho-Collados

As Artificial Intelligence (AI) is increasingly used in areas that significantly impact human lives, concerns about fairness and transparency have grown, especially regarding their impact on protected groups. Recently, the intersection of…

人工智能 · 计算机科学 2025-05-05 Vasiliki Papanikou , Danae Pla Karidi , Evaggelia Pitoura , Emmanouil Panagiotou , Eirini Ntoutsi
‹ 上一页 1 2 3 10 下一页 ›