Computation and Language · Computer Science
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
Hyukhun Koh, Dohyung Kim, Minwoo Lee, Kyomin Jung
2024-11-15
Computation and Language · Computer Science
Toxicity Inspector: A Framework to Evaluate Ground Truth in Toxicity Detection Through Feedback
Huriyyah Althunayan, Rahaf Bahlas, Manar Alharbi, Lena Alsuwailem +2
2023-05-19
Machine Learning · Computer Science
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
Sergei Berezin, Reza Farahbakhsh, Noel Crespi
2026-05-13
Computation and Language · Computer Science
Benchmark on Peer Review Toxic Detection: A Challenging Task with a New Dataset
Man Luo, Bradley Peterson, Rafael Gan, Hari Ramalingame +4
2025-02-05
Computation and Language · Computer Science
Challenges for Toxic Comment Classification: An In-Depth Error Analysis
Betty van Aken, Julian Risch, Ralf Krestel, Alexander Löser
2018-09-21
Computation and Language · Computer Science
Toxicity Detection can be Sensitive to the Conversational Context
Alexandros Xenos, John Pavlopoulos, Ion Androutsopoulos, Lucas Dixon +2
2021-11-22
Cryptography and Security · Computer Science
Toxicity Detection towards Adaptability to Changing Perturbations
Hankun Kang, Jianhao Chen, Yongqi Li, Xin Miao +4
2025-03-05
Computation and Language · Computer Science
Challenges in Detoxifying Language Models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri +6
2021-09-16
Computation and Language · Computer Science
ToxiCraft: A Novel Framework for Synthetic Generation of Harmful Information
Zheng Hui, Zhaoxiao Guo, Hang Zhao, Juanyong Duan +1
2025-04-16
Machine Learning · Computer Science
On Robustness of Linear Classifiers to Targeted Data Poisoning
Nakshatra Gupta, Sumanth Prabhu, Supratik Chakraborty, R Venkatesh
2025-11-18
Machine Learning · Computer Science
When Bad Data Leads to Good Models
Kenneth Li, Yida Chen, Fernanda Viégas, Martin Wattenberg
2025-05-09
Computation and Language · Computer Science
Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
Zhiqiang Kou, Junyang Chen, Xin-Qiang Cai, Ming-Kun Xie +7
2025-10-20
Computation and Language · Computer Science
Realistic Evaluation of Toxicity in Large Language Models
Tinh Son Luong, Thanh-Thien Le, Linh Ngo Van, Thien Huu Nguyen
2024-05-21
Computation and Language · Computer Science
Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models
Maribeth Rauh, John Mellor, Jonathan Uesato, Po-Sen Huang +8
2022-10-31
Computation and Language · Computer Science
QUACKIE: A NLP Classification Task With Ground Truth Explanations
Yves Rychener, Xavier Renard, Djamé Seddah, Pascal Frossard +1
2020-12-29
Computation and Language · Computer Science
Concept-Based Interpretability for Toxicity Detection
Samarth Garg, Divya Singh, Deeksha Varshney, Mamta
2025-12-16
Computation and Language · Computer Science
NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Oscar Sainz, Jon Ander Campos, Iker García-Ferrero, Julen Etxaniz +2
2023-10-30