English

VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark

Computation and Language 2026-01-26 v4 Machine Learning

Abstract

We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (VLMs) interpret and reason over visual and textual information beyond English. VMMU consists of 2.5k multimodal questions across 7 tasks, covering a diverse range of problem contexts, including STEM problem solving, data interpretation, rule-governed visual reasoning, and abstract visual reasoning. All questions require genuine multimodal integration, rather than reliance on text-only cues or OCR-based shortcuts. We evaluate a diverse set of state-of-the-art proprietary and open-source VLMs on VMMU. Despite strong Vietnamese OCR performance, proprietary models achieve only 66% mean accuracy. Further analysis shows that the primary source of failure is not OCR, but instead multimodal grounding and reasoning over text and visual evidence. Code and data are available at https://vmmu-bench.github.io/

Keywords

Cite

@article{arxiv.2508.13680,
  title  = {VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark},
  author = {Vy Tuong Dang and An Vo and Emilio Villa-Cueva and Quang Tau and Duc Dm and Thamar Solorio and Daeyoung Kim},
  journal= {arXiv preprint arXiv:2508.13680},
  year   = {2026}
}
R2 v1 2026-07-01T04:56:27.195Z