English

MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents

Artificial Intelligence 2026-04-29 v2

Abstract

Large Vision-Language Models (LVLMs) have shown strong potential as multilingual Graphical User Interface (GUI) agents, as evidenced by existing GUI benchmarks. However, these benchmarks exhibit two primary limitations: (1) although Perception and Reasoning (P&R) capabilities are fundamental for GUI agents, current benchmarks lack fine-grained diagnostics to identify which specific capabilities lead to task failures, hindering targeted improvements; (2) existing benchmarks fail to provide a strictly aligned cross-lingual evaluation environment, introducing confounding factors that prevent isolating the language impact on GUI agent performance. To address these issues, we propose the Multilingual P&R GUI Benchmark (MPR-GUI-Bench), featuring strictly aligned environments across six languages and eight fine-grained P&R tasks. Our benchmark reveals consistent P&R gaps between English and non-English settings, particularly on reasoning-intensive tasks. To leverage the superior English P&R capabilities for bridging cross-lingual gaps, we identify layers sensitive to language and propose GUI-XLI, a GUI Cross-Lingual Intervention method that aligns non-English hidden states with their English counterparts at these layers during inference. Experiments show that GUI-XLI effectively reduces the cross-lingual gaps, with an average gain of 6.5% in non-English settings.

Keywords

Cite

@article{arxiv.2512.00756,
  title  = {MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents},
  author = {Ruihan Chen and Qiming Li and Xiaocheng Feng and Weihong Zhong and Xiaoliang Yang and Yuxuan Gu and Zekun Zhou and Yunfei Lu and Haoyu Ren and Kun Chen and Dandan Tu and Bing Qin},
  journal= {arXiv preprint arXiv:2512.00756},
  year   = {2026}
}

Comments

35pages, 15figures

R2 v1 2026-07-01T08:01:29.722Z