English

Probing Association Biases in LLM Moderation Over-Sensitivity

Computation and Language 2026-03-19 v2 Artificial Intelligence

Abstract

Large Language Models are widely used for content moderation but often present certain over-sensitivity, leading to misclassification of benign content and rejecting safe user commands. While previous research attributes this issue primarily to the presence of explicit offensive triggers, we statistically reveal a deeper connection beyond token level: When behaving over-sensitively, particularly on decontextualized statements, LLMs exhibit systematic topic-toxicity association patterns that go beyond explicit offensive triggers. To characterize these patterns, we propose Topic Association Analysis, a behavior-based probe that elicits short contextual scenarios for benign inputs and quantifies topic amplification between the scenario and the original comment. Across multiple LLMs and large-scale data, we find that more advanced models (e.g., GPT-4 Turbo) show stronger topic-association skew in false-positive cases despite lower overall false-positive rates. Moreover, via controlled prefix interventions, we show that topic cues can measurably shift false-positive rates, indicating that topic framing is decision-relevant. These results suggest that mitigating over-sensitivity may require addressing learned topic associations in addition to keyword-based filtering.

Keywords

Cite

@article{arxiv.2505.23914,
  title  = {Probing Association Biases in LLM Moderation Over-Sensitivity},
  author = {Yuxin Wang and Botao Yu and Ivory Yang and Saeed Hassanpour and Soroush Vosoughi},
  journal= {arXiv preprint arXiv:2505.23914},
  year   = {2026}
}

Comments

Preprint

R2 v1 2026-07-01T02:49:17.563Z