Computation and Language · Computer Science
Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
Rijul Magu, Arka Dutta, Sean Kim, Ashiqur R. KhudaBukhsh +1
2026-01-30
Computation and Language · Computer Science
Toxicity Inspector: A Framework to Evaluate Ground Truth in Toxicity Detection Through Feedback
Huriyyah Althunayan, Rahaf Bahlas, Manar Alharbi, Lena Alsuwailem +2
2023-05-19
Computation and Language · Computer Science
Ethical and social risks of harm from Language Models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin +19
2021-12-09
Computation and Language · Computer Science
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
Zhiqiang Kou, Junyang Chen, Xin-Qiang Cai, Ming-Kun Xie +7
2025-10-20
Computation and Language · Computer Science
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
Luiza Pozzobon, Patrick Lewis, Sara Hooker, Beyza Ermis
2024-05-31
Computation and Language · Computer Science
Mitigating Racial Biases in Toxic Language Detection with an Equity-Based Ensemble Framework
Matan Halevy, Camille Harris, Amy Bruckman, Diyi Yang +1
2021-09-28
Computation and Language · Computer Science
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
Atharvan Dogra, Soumya Suvra Ghosal, Ameet Deshpande, Ashwin Kalyan +1
2025-10-22
Computation and Language · Computer Science
Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models
Maribeth Rauh, John Mellor, Jonathan Uesato, Po-Sen Huang +8
2022-10-31
Social and Information Networks · Computer Science
Down the Rabbit Hole: Detecting Online Extremism, Radicalisation, and Politicised Hate Speech
Jarod Govers, Philip Feldman, Aaron Dant, Panos Patros
2023-01-30
Computation and Language · Computer Science
ToXCL: A Unified Framework for Toxic Speech Detection and Explanation
Nhat M. Hoang, Xuan Long Do, Duc Anh Do, Duc Anh Vu +1
2024-05-21
Computation and Language · Computer Science
Probing LLMs for hate speech detection: strengths and vulnerabilities
Sarthak Roy, Ashish Harshavardhan, Animesh Mukherjee, Punyajoy Saha
2023-10-31
Computation and Language · Computer Science
Mitigating Biases in Toxic Language Detection through Invariant Rationalization
Yung-Sung Chuang, Mingye Gao, Hongyin Luo, James Glass +3
2021-06-15
Computation and Language · Computer Science
Unveiling the Implicit Toxicity in Large Language Models
Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang +3
2023-11-30
Computation and Language · Computer Science
Defining, Understanding, and Detecting Online Toxicity: Challenges and Machine Learning Approaches
Gautam Kishore Shahi, Tim A. Majchrzak
2025-09-19
Computation and Language · Computer Science
ROBBIE: Robust Bias Evaluation of Large Generative Language Models
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung +6
2023-12-01
Computers and Society · Computer Science
Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
Smita Khapre, Melkamu Abay Mersha, Hassan Shakil, Jonali Baruah +1
2025-10-01
Computation and Language · Computer Science
A Framework to Assess Multilingual Vulnerabilities of LLMs
Likai Tang, Niruth Bogahawatta, Yasod Ginige, Jiarui Xu +3
2025-03-18
Computation and Language · Computer Science
Unveiling Safety Vulnerabilities of Large Language Models
George Kour, Marcel Zalmanovici, Naama Zwerdling, Esther Goldbraich +4
2023-11-08
Software Engineering · Computer Science
How Toxic Can You Get? Search-based Toxicity Testing for Large Language Models
Simone Corbo, Luca Bancale, Valeria De Gennaro, Livia Lestingi +2
2026-02-06
Computation and Language · Computer Science
RECAST: Interactive Auditing of Automatic Toxicity Detection Models
Austin P. Wright, Omar Shaikh, Haekyu Park, Will Epperson +4
2020-07-02
Computation and Language · Computer Science
Measuring Misogyny in Natural Language Generation: Preliminary Results from a Case Study on two Reddit Communities
Aaron J. Snoswell, Lucinda Nelson, Hao Xue, Flora D. Salim +2
2023-12-07
Computation and Language · Computer Science
Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
Muhammed Saeed, Muhammad Abdul-mageed, Shady Shehata
2026-03-31
Computation and Language · Computer Science
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
Ahmed Sharshar, Hosam Elgendy, Saad El Dine Ahmed, Yasser Rohaim +1
2026-03-20