English

Towards A Friendly Online Community: An Unsupervised Style Transfer Framework for Profanity Redaction

Computation and Language 2020-11-03 v1 Machine Learning Social and Information Networks

Abstract

Offensive and abusive language is a pressing problem on social media platforms. In this work, we propose a method for transforming offensive comments, statements containing profanity or offensive language, into non-offensive ones. We design a RETRIEVE, GENERATE and EDIT unsupervised style transfer pipeline to redact the offensive comments in a word-restricted manner while maintaining a high level of fluency and preserving the content of the original text. We extensively evaluate our method's performance and compare it to previous style transfer models using both automatic metrics and human evaluations. Experimental results show that our method outperforms other models on human evaluations and is the only approach that consistently performs well on all automatic evaluation metrics.

Keywords

Cite

@article{arxiv.2011.00403,
  title  = {Towards A Friendly Online Community: An Unsupervised Style Transfer Framework for Profanity Redaction},
  author = {Minh Tran and Yipeng Zhang and Mohammad Soleymani},
  journal= {arXiv preprint arXiv:2011.00403},
  year   = {2020}
}

Comments

COLING 2020

R2 v1 2026-06-23T19:48:53.097Z