English

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

Sound 2026-01-27 v1 Computation and Language Computer Vision and Pattern Recognition Multimedia Audio and Speech Processing

Abstract

Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we introduce AVMeme Exam, a human-curated benchmark of over one thousand iconic Internet sounds and videos spanning speech, songs, music, and sound effects. Each meme is paired with a unique Q&A assessing levels of understanding from surface content to context and emotion to usage and world knowledge, along with metadata such as original year, transcript, summary, and sensitivity. We systematically evaluate state-of-the-art multimodal large language models (MLLMs) alongside human participants using this benchmark. Our results reveal a consistent limitation: current models perform poorly on textless music and sound effects, and struggle to think in context and in culture compared to surface content. These findings highlight a key gap in human-aligned multimodal intelligence and call for models that can perceive contextually and culturally beyond the surface of what they hear and see. Project page: avmemeexam.github.io/public

Keywords

Cite

@article{arxiv.2601.17645,
  title  = {AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking},
  author = {Xilin Jiang and Qiaolin Wang and Junkai Wu and Xiaomin He and Zhongweiyang Xu and Yinghao Ma and Minshuo Piao and Kaiyi Yang and Xiuwen Zheng and Riki Shimizu and Yicong Chen and Arsalan Firoozi and Gavin Mischler and Sukru Samet Dindar and Richard Antonello and Linyang He and Tsun-An Hsieh and Xulin Fan and Yulun Wu and Yuesheng Ma and Chaitanya Amballa and Weixiong Chen and Jiarui Hai and Ruisi Li and Vishal Choudhari and Cong Han and Yinghao Aaron Li and Adeen Flinker and Mounya Elhilali and Emmanouil Benetos and Mark Hasegawa-Johnson and Romit Roy Choudhury and Nima Mesgarani},
  journal= {arXiv preprint arXiv:2601.17645},
  year   = {2026}
}

Comments

avmemeexam.github.io/public