English

LifeTox: Unveiling Implicit Toxicity in Life Advice

Computation and Language 2024-03-20 v2

Abstract

As large language models become increasingly integrated into daily life, detecting implicit toxicity across diverse contexts is crucial. To this end, we introduce LifeTox, a dataset designed for identifying implicit toxicity within a broad range of advice-seeking scenarios. Unlike existing safety datasets, LifeTox comprises diverse contexts derived from personal experiences through open-ended questions. Experiments demonstrate that RoBERTa fine-tuned on LifeTox matches or surpasses the zero-shot performance of large language models in toxicity classification tasks. These results underscore the efficacy of LifeTox in addressing the complex challenges inherent in implicit toxicity. We open-sourced the dataset\footnote{\url{https://huggingface.co/datasets/mbkim/LifeTox}} and the LifeTox moderator family; 350M, 7B, and 13B.

Keywords

Cite

@article{arxiv.2311.09585,
  title  = {LifeTox: Unveiling Implicit Toxicity in Life Advice},
  author = {Minbeom Kim and Jahyun Koo and Hwanhee Lee and Joonsuk Park and Hwaran Lee and Kyomin Jung},
  journal= {arXiv preprint arXiv:2311.09585},
  year   = {2024}
}

Comments

11 pages, 5 figures, NAACL 2024

R2 v1 2026-06-28T13:22:58.252Z