English

Sign Language: Towards Sign Understanding for Robot Autonomy

Robotics 2025-09-17 v2

Abstract

Navigational signs are common aids for human wayfinding and scene understanding, but are underutilized by robots. We argue that they benefit robot navigation and scene understanding, by directly encoding privileged information on actions, spatial regions, and relations. Interpreting signs in open-world settings remains a challenge owing to the complexity of scenes and signs, but recent advances in vision-language models (VLMs) make this feasible. To advance progress in this area, we introduce the task of navigational sign understanding which parses locations and associated directions from signs. We offer a benchmark for this task, proposing appropriate evaluation metrics and curating a test set capturing signs with varying complexity and design across diverse public spaces, from hospitals to shopping malls to transport hubs. We also provide a baseline approach using VLMs, and demonstrate their promise on navigational sign understanding. Code and dataset are available on Github.

Keywords

Cite

@article{arxiv.2506.02556,
  title  = {Sign Language: Towards Sign Understanding for Robot Autonomy},
  author = {Ayush Agrawal and Joel Loo and Nicky Zimmerman and David Hsu},
  journal= {arXiv preprint arXiv:2506.02556},
  year   = {2025}
}

Comments

This work has been submitted to the IEEE for possible publication

R2 v1 2026-07-01T02:56:10.622Z