English

IPS: In-Prompt Process Supervision for Short Video Content Moderation

Computation and Language 2026-05-05 v3 Artificial Intelligence

Abstract

Multimodal large language models (MLLMs) are effective at capturing the semantics of short video content; however, they often fail to attend to the policy-specific details required for reliable content moderation. To address this limitation, we introduce IPS, a novel framework that integrates In-prompt Process Supervision into MLLMs by introducing sequential reasoning over ancillary questions during fine-tuning. IPS consistently outperforms baseline MLLMs on public and proprietary benchmarks. Moreover, replacing human-annotated ancillary labels with MLLM-generated ones results in only marginal performance degradation, demonstrating robustness to noisy supervision and strong scalability with model-generated annotations. These findings establish IPS as a scalable and effective solution for complex multimodal classification in large-scale industrial settings.

Keywords

Cite

@article{arxiv.2412.15251,
  title  = {IPS: In-Prompt Process Supervision for Short Video Content Moderation},
  author = {Mingchao Liu and Yu Sun and Ruixiao Sun and Xin Dong and Xiang Shen and Hongwei Wang and Hongyu Xiong and Yang Song},
  journal= {arXiv preprint arXiv:2412.15251},
  year   = {2026}
}

Comments

7 pages(excluding reference and appendix), 8 figures

R2 v1 2026-06-28T20:42:52.298Z