English

RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos

Computer Vision and Pattern Recognition 2025-10-13 v1 Artificial Intelligence

Abstract

Recently, Multi-modal Large Language Models (MLLMs) have demonstrated significant performance across various video understanding tasks. However, their robustness, particularly when faced with manipulated video content, remains largely unexplored. In this paper, we introduce Ro-Bench, the first benchmark for evaluating MLLMs on dynamic out-of-distribution (OOD) counterfactual video test sets. Ro-Bench incorporates high-quality, diverse and temporally relevant video data, by editing Style, Object, Background and their compositions. We evaluated eight recent video MLLMs and found that current models exhibit substantial performance degradation on Ro-Bench when exposed to counterfactual video content. Furthermore, we demonstrate that fine-tuning MLLMs with counterfactual data enhances robustness, achieving a 21.73% performance increase on Ro-Bench and a 12.78% improvement across 20 tasks in the MVBench dataset. These findings underscore the effectiveness of counterfactual data in enhancing the video understanding ability of MLLMs. The code and data will be released shortly.

Keywords

Cite

@article{arxiv.2510.08936,
  title  = {RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos},
  author = {Zixi Yang and Jiapeng Li and Muxi Diao and Yinuo Jing and Kongming Liang},
  journal= {arXiv preprint arXiv:2510.08936},
  year   = {2025}
}