English

Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models

Computation and Language 2022-10-31 v2 Artificial Intelligence Computers and Society

Abstract

Large language models produce human-like text that drive a growing number of applications. However, recent literature and, increasingly, real world observations, have demonstrated that these models can generate language that is toxic, biased, untruthful or otherwise harmful. Though work to evaluate language model harms is under way, translating foresight about which harms may arise into rigorous benchmarks is not straightforward. To facilitate this translation, we outline six ways of characterizing harmful text which merit explicit consideration when designing new benchmarks. We then use these characteristics as a lens to identify trends and gaps in existing benchmarks. Finally, we apply them in a case study of the Perspective API, a toxicity classifier that is widely used in harm benchmarks. Our characteristics provide one piece of the bridge that translates between foresight and effective evaluation.

Keywords

Cite

@article{arxiv.2206.08325,
  title  = {Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models},
  author = {Maribeth Rauh and John Mellor and Jonathan Uesato and Po-Sen Huang and Johannes Welbl and Laura Weidinger and Sumanth Dathathri and Amelia Glaese and Geoffrey Irving and Iason Gabriel and William Isaac and Lisa Anne Hendricks},
  journal= {arXiv preprint arXiv:2206.08325},
  year   = {2022}
}

Comments

Accepted to NeurIPS 2022 Datasets and Benchmarks Track; 10 pages plus appendix