English

From Language to Logic: A Theoretical Architecture for VLM-Grounded Safe Navigation

Robotics 2026-05-07 v1

Abstract

We propose an architecture for integrating high-level, human-provided safety rules and operator-aligned semantic preferences into autonomous robot navigation in unstructured outdoor environments. In our approach, natural-language rules are translated into Signal Temporal Logic (STL) specifications that guide planning and navigation during runtime. Persistent, environment-centric rules and terrain preferences are grounded into a 2D cost map, while temporally dynamic requirements are expressed as STL specifications to be monitored during runtime. We hypothesize the use of Vision-Language Models (VLMs) for zero-shot scene understanding, enabling mapping between human instructions, semantic features, and environmental constraints. Within this framework, we construct an illustrative navigation model that is designed to satisfy a set of STL-encoded specifications and soft operator preferences through formal satisfaction metrics embedded into environmental properties and runtime monitoring.

Keywords

Cite

@article{arxiv.2605.04327,
  title  = {From Language to Logic: A Theoretical Architecture for VLM-Grounded Safe Navigation},
  author = {Kristy Sakano and Kalonji Harrington and Mumu Xu},
  journal= {arXiv preprint arXiv:2605.04327},
  year   = {2026}
}

Comments

8 pages, 3 figures, to be published in ICUAS 2026 conference proceedings

R2 v1 2026-07-01T12:51:54.421Z