English

Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework

Computers and Society 2026-05-05 v1 Computation and Language

Abstract

The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy. Traditional content moderation systems rely on centralised, top-down rules, often failing to accommodate the subjective nature of harm perception. This paper proposes an LLM-based multi-agent personalised inference framework that filters content based on unique sensitivity profiles of individual users. Our architecture combines domain-specific Expert Agents, a Manager Agent for orchestrating content analysis and agent selection, and a Ghost Profile Agent for simulating user perspectives, to inform moderation decisions. Evaluated against a range of non-personalised baselines, the system demonstrates up to a 32% improvement in accuracy, showing increased alignment with individual user sensitivities. Beyond technical performance, our framework provides policy-relevant insights for platform governance, providing a scalable way to reconcile moderation policies with societal and individual digital rights

Keywords

Cite

@article{arxiv.2605.01416,
  title  = {Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework},
  author = {Ewelina Gajewska and Michal Wawer and Katarzyna Budzynska and Jaroslaw A. Chudziak},
  journal= {arXiv preprint arXiv:2605.01416},
  year   = {2026}
}

Comments

The paper has been accepted to the 34th European Conference on Information Systems (ECIS 2026). The official paper version will appear in the conference proceedings

R2 v1 2026-07-01T12:46:38.767Z