English

Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness

Computer Vision and Pattern Recognition 2026-03-03 v2 Artificial Intelligence

Abstract

Cultural awareness capabilities have emerged as a critical capability for Multimodal Large Language Models (MLLMs). However, current benchmarks lack progressed difficulty in their task design and are deficient in cross-lingual tasks. Moreover, current benchmarks often use real-world images. Each real-world image typically contains one culture, making these benchmarks relatively easy for MLLMs. Based on this, we propose C3^3B (Comics Cross-Cultural Benchmark), a novel multicultural, multitask and multilingual cultural awareness capabilities benchmark. C3^3B comprises over 2000 images and over 18000 QA pairs, constructed on three tasks with progressed difficulties, from basic visual recognition to higher-level cultural conflict understanding, and finally to cultural content generation. We conducted evaluations on 11 open-source MLLMs, revealing a significant performance gap between MLLMs and human performance. The gap demonstrates that C3^3B poses substantial challenges for current MLLMs, encouraging future research to advance the cultural awareness capabilities of MLLMs.

Keywords

Cite

@article{arxiv.2510.00041,
  title  = {Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness},
  author = {Yuchen Song and Andong Chen and Wenxin Zhu and Kehai Chen and Xuefeng Bai and Muyun Yang and Tiejun Zhao},
  journal= {arXiv preprint arXiv:2510.00041},
  year   = {2026}
}