English

Shieldstral

Computation and Language 2026-07-28 v1 Computer Vision and Pattern Recognition

Abstract

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7×\times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.

Cite

@article{arxiv.2607.25857,
  title  = {Shieldstral},
  author = {Antonia Calvi and Avinash Sooriyarachchi and Giada Pistilli and Guillaume Lample and Maarten Buyl and Maximilian Augustin and Maximilian Müller and Pierre Stock and Tom Bewley and Wassim Bouaziz and Yimu Pan},
  journal= {arXiv preprint arXiv:2607.25857},
  year   = {2026}
}