English
Related papers

Related papers: Waypoint Models for Instruction-guided Navigation …

200 papers

Vision-language models (VLMs) achieve strong performance on multimodal benchmarks, but may still lack robust control over basic visual operations. We study \textit{line tracing}, where a model must follow a selected visual path through…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Hyesoo Hong , Minsoo Kim , Wonje Jeung , Sangyeon Yoon , Dongjae Jeon , Albert No

Vision-and-Language Navigation (VLN) is a cornerstone of embodied intelligence. However, current agents often suffer from significant performance degradation when transitioning from simulation to real-world deployment, primarily due to…

Vision-and-language navigation (VLN) aims to enable embodied agents to navigate in realistic environments using natural language instructions. Given the scarcity of domain-specific training data and the high diversity of image and language…

Computer Vision and Pattern Recognition · Computer Science 2021-08-23 Pierre-Louis Guhur , Makarand Tapaswi , Shizhe Chen , Ivan Laptev , Cordelia Schmid

Natural language instructions for visual navigation often use scene descriptions (e.g., "bedroom") and object references (e.g., "green chairs") to provide a breadcrumb trail to a goal location. This work presents a transformer-based…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Abhinav Moudgil , Arjun Majumdar , Harsh Agrawal , Stefan Lee , Dhruv Batra

Vision-and-language navigation (VLN) tasks require agents to navigate three-dimensional environments guided by natural language instructions, offering substantial potential for diverse applications. However, the scarcity of training data…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Sen Wang , Dongliang Zhou , Liang Xie , Chao Xu , Ye Yan , Erwei Yin

Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense and spatial reasoning for efficient exploration in…

Learning-based behavior prediction methods are increasingly being deployed in real-world autonomous systems, e.g., in fleets of self-driving vehicles, which are beginning to commercially operate in major cities across the world. Despite…

Machine Learning · Computer Science 2023-05-24 Boris Ivanovic , James Harrison , Marco Pavone

In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning. However, existing paradigms fall short of jointly modeling these…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Binjie Zhang , Mike Zheng Shou

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. At each navigation step, the agent selects from possible candidate locations and then…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Zihan Wang , Xiangyang Li , Jiahao Yang , Yeqi Liu , Junjie Hu , Ming Jiang , Shuqiang Jiang

We discuss the process of building semantic maps, how to interactively label entities in them, and how to use them to enable context-aware navigation behaviors in human environments. We utilize planar surfaces, such as walls and tables, and…

Robotics · Computer Science 2018-08-15 Akansel Cosgun , Henrik Christensen

Vision-and-Language Navigation (VLN) aims to enable an embodied agent to follow natural-language instructions and navigate to a target location in unseen 3D environments. We argue that adapting VLMs to VLN requires endowing them with two…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Pengna Li , Kangyi Wu , Shaoqing Xu , Fang Li , Hanbing Li , Lin Zhao , Kailin Lyu , Long Chen , Zhi-Xin Yang , Nanning Zheng

We investigate the task of object goal navigation in unknown environments where the target is specified by a semantic label (e.g. find a couch). Such a navigation task is especially challenging as it requires understanding of semantic…

Robotics · Computer Science 2022-10-18 Yash Goel , Narunas Vaskevicius , Luigi Palmieri , Nived Chebrolu , Cyrill Stachniss

In visual semantic navigation, the robot navigates to a target object with egocentric visual observations and the class label of the target is given. It is a meaningful task inspiring a surge of relevant research. However, most of the…

Artificial Intelligence · Computer Science 2021-09-21 Xinzhu Liu , Di Guo , Huaping Liu , Fuchun Sun

Building on the unprecedented capabilities of large language models for command understanding and zero-shot recognition of multi-modal vision-language transformers, visual language navigation (VLN) has emerged as an effective way to address…

Robotics · Computer Science 2024-07-11 Chashi Mahiul Islam , Shaeke Salman , Montasir Shams , Xiuwen Liu , Piyush Kumar

The performance of learned robot visuomotor policies is heavily dependent on the size and quality of the training dataset. Although large-scale robot and human datasets are increasingly available, embodiment gaps and mismatched action…

Robotics · Computer Science 2026-03-24 Yiqi Wang , Mrinal Verghese , Jeff Schneider

Identifying an appropriate task space that simplifies control solutions is important for solving robotic manipulation problems. One approach to this problem is learning an appropriate low-dimensional action space. Linear and nonlinear…

In this paper, we propose a novel approach for agent motion prediction in cluttered environments. One of the main challenges in predicting agent motion is accounting for location and context-specific information. Our main contribution is…

Robotics · Computer Science 2020-07-08 Igor Gilitschenski , Guy Rosman , Arjun Gupta , Sertac Karaman , Daniela Rus

Object-goal navigation (Object-nav) entails searching, recognizing and navigating to a target object. Object-nav has been extensively studied by the Embodied-AI community, but most solutions are often restricted to considering static…

Vision-Language-Action (VLA) models have demonstrated strong performance across a wide range of robotic manipulation tasks. Despite the success, extending large pretrained Vision-Language Models (VLMs) to the action space can induce…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Yiye Chen , Yanan Jian , Xiaoyi Dong , Shuxin Cao , Jing Wu , Patricio Vela , Benjamin E. Lundell , Dongdong Chen

Service robots are increasingly deployed in diverse and dynamic environments, where both physical layouts and social contexts change over time and across locations. In these unstructured settings, conventional navigation systems that rely…

Robotics · Computer Science 2025-07-16 Yanbo Wang , Zipeng Fang , Lei Zhao , Weidong Chen