English

MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments

Distributed, Parallel, and Cluster Computing 2026-02-09 v1

Abstract

Current mobile GUI agent benchmarks systematically fail to assess memory capabilities, with only 5.2-11.8% memory-related tasks and no cross-session learning evaluation. We introduce MemGUI-Bench, a comprehensive memory-centric benchmark with pass@k and staged LLM-as-judge evaluation. Our contributions include: (1) a systematic memory taxonomy analyzing 11 agents across 5 architectures; (2) 128 tasks across 26 applications where 89.8% challenge memory through cross-temporal and cross-spatial retention; (3) MemGUI-Eval, an automated pipeline with Progressive Scrutiny and 7 hierarchical metrics; and (4) RQ-driven assessment of 11 state-of-the-art agents. Our experiments reveal significant memory deficits across all evaluated systems, identify 5 distinct failure modes, and synthesize 5 actionable design implications. All resources including code, benchmark, and evaluation results will be \textbf{\textit{fully open-sourced and continuously maintained}} at https://lgy0404.github.io/MemGUI-Bench/.

Keywords

Cite

@article{arxiv.2602.06075,
  title  = {MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments},
  author = {Guangyi Liu and Pengxiang Zhao and Yaozhen Liang and Qinyi Luo and Shunye Tang and Yuxiang Chai and Weifeng Lin and Han Xiao and WenHao Wang and Siheng Chen and Zhengxi Lu and Gao Wu and Hao Wang and Liang Liu and Yong Liu},
  journal= {arXiv preprint arXiv:2602.06075},
  year   = {2026}
}

Comments

https://lgy0404.github.io/MemGUI-Bench/

R2 v1 2026-07-01T10:23:12.688Z