English

Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification

Computation and Language 2025-10-24 v3 Artificial Intelligence

Abstract

As large language models (LLMs) become increasingly prevalent in global applications, ensuring that they are toxicity-free across diverse linguistic contexts remains a critical challenge. We explore "Cross-lingual Detoxification", a cross-lingual paradigm that mitigates toxicity, enabling detoxification capabilities to transfer between high and low-resource languages across different script families. We analyze cross-lingual detoxification's effectiveness through 392 extensive settings to evaluate toxicity reduction in cross-distribution settings with limited data and investigate how mitigation impacts model performance on non-toxic tasks, revealing trade-offs between safety and knowledge preservation. Our code and dataset are publicly available at https://github.com/himanshubeniwal/Breaking-mBad.

Keywords

Cite

@article{arxiv.2505.16722,
  title  = {Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification},
  author = {Himanshu Beniwal and Youngwoo Kim and Maarten Sap and Soham Dan and Thomas Hartvigsen},
  journal= {arXiv preprint arXiv:2505.16722},
  year   = {2025}
}

Comments

Accepted at MELT Workshop @ COLM 2025

R2 v1 2026-07-01T02:31:41.656Z