SEE: Structure-aware Exploring \& Exploiting for Long-horizon GUI Agent Trajectory Synthesis
Abstract
Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-horizon interaction trajectories collected from element-rich and rapidly evolving apps. Existing pipelines often rely on costly human demonstrations or on-policy framework, which tends to over-sample common flows while missing rare transitions and complex multi-step procedures. To address this problem, we propose SEE, a two-stage data synthesis framework consisting of (i) an efficient exploration stage that builds an explicit UI transition graph over screens and elements, and (ii) a graph-based synthesis stage that composes diverse multi-step trajectories via planning and controlled sampling. This design yields reproducible and explainable data generation, while explicitly preventing spurious cycles and enabling long-horizon composition. Across multiple real-world apps, SEE produces trajectories with an average length of 14.8 steps while avoiding spurious loops, and agents fine-tuned on SEE achieve improved task success and generalization to unseen screens. We will publicly release our synthesis code and dataset.
Cite
@article{arxiv.2607.18046,
title = {SEE: Structure-aware Exploring \& Exploiting for Long-horizon GUI Agent Trajectory Synthesis},
author = {Zhuohang Fan and Beichen Zhang and Yuanfa Li and Changqiao Wu and Wei Liu and Jian Luan and Weigang Zhang},
journal= {arXiv preprint arXiv:2607.18046},
year = {2026}
}
Comments
Accepted by ACM International Conference on Multimedia 2026 (ACM MM 2026)