English

From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces

Machine Learning 2023-12-08 v2 Computation and Language Computer Vision and Pattern Recognition Human-Computer Interaction

Abstract

Much of the previous work towards digital agents for graphical user interfaces (GUIs) has relied on text-based representations (derived from HTML or other structured data sources), which are not always readily available. These input representations have been often coupled with custom, task-specific action spaces. This paper focuses on creating agents that interact with the digital world using the same conceptual interface that humans commonly use -- via pixel-based screenshots and a generic action space corresponding to keyboard and mouse actions. Building upon recent progress in pixel-based pretraining, we show, for the first time, that it is possible for such agents to outperform human crowdworkers on the MiniWob++ benchmark of GUI-based instruction following tasks.

Keywords

Cite

@article{arxiv.2306.00245,
  title  = {From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces},
  author = {Peter Shaw and Mandar Joshi and James Cohan and Jonathan Berant and Panupong Pasupat and Hexiang Hu and Urvashi Khandelwal and Kenton Lee and Kristina Toutanova},
  journal= {arXiv preprint arXiv:2306.00245},
  year   = {2023}
}
R2 v1 2026-06-28T10:52:43.138Z