English

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models

Computer Vision and Pattern Recognition 2026-05-26 v2 Artificial Intelligence

Abstract

Modern Large Multimodal Models (LMMs) have demonstrated extraordinary ability in static image and single-state spatial-temporal understanding. However, their capacity to comprehend the dynamic changes of objects within a shared spatial context between two distinct video observations, remains largely unexplored. This ability to reason about transformations within a consistent environment is particularly crucial for advancements in the field of spatial intelligence. In this paper, we introduce M3VerseM^3-Verse, a Multi-Modal, Multi-State, Multi-Dimensional benchmark, to formally evaluate this capability. It is built upon paired videos that provide multi-perspective observations of an indoor scene before and after a state change. The benchmark contains a total of 270 scenes and 2,932 questions, which are categorized into over 50 subtasks that probe 4 core capabilities. We evaluate 16 state-of-the-art LMMs and observe their limitations in tracking state transitions. To address these challenges, we further propose a simple yet effective baseline that achieves significant performance improvements in multi-state perception. M3VerseM^3-Verse thus provides a challenging new testbed to catalyze the development of next-generation models with a more holistic understanding of our dynamic visual world. You can get the construction pipeline from https://github.com/Wal-K-aWay/M3-Verse_pipeline and full benchmark data from https://www.modelscope.cn/datasets/WalKaWay/M3-Verse.

Keywords

Cite

@article{arxiv.2512.18735,
  title  = {$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models},
  author = {Kewei Wei and Bocheng Hu and Jie Cao and Xiaohan Chen and Zhengxi Lu and Wubing Xia and Weili Xu and Jiaao Wu and Junchen He and Mingyu Jia and Ciyun Zhao and Ye Sun and Yizhi Li and Zhonghan Zhao and Jian Zhang and Gaoang Wang},
  journal= {arXiv preprint arXiv:2512.18735},
  year   = {2026}
}
R2 v1 2026-07-01T08:35:32.953Z