Computation and Language · Computer Science
Incorporating Human Explanations for Robust Hate Speech Detection
Jennifer L. Chen, Faisal Ladhak, Daniel Li, Noémie Elhadad
2024-11-12
Computation and Language · Computer Science
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
Tharindu Kumarage, Amrita Bhattacharjee, Joshua Garland
2024-03-14
Computers and Society · Computer Science
Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception
Luyang Lin, Lingzhi Wang, Jinsong Guo, Kam-Fai Wong
2024-12-11
Machine Learning · Computer Science
A Comprehensive Study of Implicit and Explicit Biases in Large Language Models
Fatima Kazi, Alex Young, Yash Inani, Setareh Rafatirad
2025-11-19
Computation and Language · Computer Science
Towards Understanding and Mitigating Social Biases in Language Models
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan Salakhutdinov
2021-06-25
Computation and Language · Computer Science
Improving Counterfactual Generation for Fair Hate Speech Detection
Aida Mostafazadeh Davani, Ali Omrani, Brendan Kennedy, Mohammad Atari +2
2021-08-05
Computation and Language · Computer Science
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
Jisu Shin, Hoyun Song, Huije Lee, Soyeong Jeong +1
2024-06-07
Computation and Language · Computer Science
Sociodemographic Bias in Language Models: A Survey and Forward Path
Vipul Gupta, Pranav Narayanan Venkit, Shomir Wilson, Rebecca J. Passonneau
2024-08-15
Computation and Language · Computer Science
LLMs and Finetuning: Benchmarking cross-domain performance for hate speech detection
Ahmad Nasir, Aadish Sharma, Kokil Jaidka, Saifuddin Ahmed
2025-05-01
Cryptography and Security · Computer Science
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
Xinyue Shen, Yixin Wu, Yiting Qu, Michael Backes +2
2025-01-29
Computation and Language · Computer Science
Personalisation or Prejudice? Addressing Geographic Bias in Hate Speech Detection using Debias Tuning in Large Language Models
Paloma Piot, Patricia Martín-Rodilla, Javier Parapar
2025-05-06
Computation and Language · Computer Science
Who Speaks Matters: Analysing the Influence of the Speaker's Ethnicity on Hate Classification
Ananya Malik, Kartik Sharma, Shaily Bhatt, Lynnette Hui Xian Ng
2025-10-14
Computation and Language · Computer Science
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
Shachi H Kumar, Saurav Sahay, Sahisnu Mazumder, Eda Okur +5
2024-08-08
Computation and Language · Computer Science
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
Ayan Majumdar, Feihao Chen, Jinghui Li, Xiaozhen Wang
2026-04-10
Computers and Society · Computer Science
An Investigation of Large Language Models for Real-World Hate Speech Detection
Keyan Guo, Alexander Hu, Jaden Mu, Ziheng Shi +3
2024-01-09
Computation and Language · Computer Science
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
Sanjeeevan Selvaganapathy, Mehwish Nasim
2026-05-05
Computation and Language · Computer Science
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
Charaka Vinayak Kumar, Ashok Urlana, Gopichand Kanumolu, Bala Mallikarjunarao Garlapati +1
2025-05-28
Computation and Language · Computer Science
Measuring Stereotype and Deviation Biases in Large Language Models
Daniel Wang, Eli Brignac, Minjia Mao, Xiao Fang
2026-05-20