English
Related papers

Related papers: NavGPT: Explicit Reasoning in Vision-and-Language …

200 papers

Visual navigation in unknown environments based solely on natural language descriptions is a key capability for intelligent robots. In this work, we propose a navigation framework built upon off-the-shelf Visual Language Models (VLMs),…

Robotics · Computer Science 2025-08-08 Weifan Zhang , Tingguang Li , Yuzhen Liu

The Zero-Shot Object Navigation (ZSON) task requires embodied agents to find a previously unseen object by navigating in unfamiliar environments. Such a goal-oriented exploration heavily relies on the ability to perceive, understand, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Linqing Zhong , Chen Gao , Zihan Ding , Yue Liao , Huimin Ma , Shifeng Zhang , Xu Zhou , Si Liu

Vision-and-Language Navigation (VLN) requires an embodied agent to ground complex natural-language instructions into long-horizon navigation in unseen environments. While Vision-Language Models (VLMs) offer strong 2D semantic understanding,…

Robotics · Computer Science 2026-03-19 Zihao Xin , Wentong Li , Yixuan Jiang , Ziyuan Huang , Bin Wang , Piji Li , Jianke Zhu , Jie Qin , Shengjun Huang

VLA models have shown promising potential in embodied navigation by unifying perception and planning while inheriting the strong generalization abilities of large VLMs. However, most existing VLA models rely on reactive mappings directly…

Robotics · Computer Science 2026-01-14 Shaoan Wang , Yuanfei Luo , Xingyu Chen , Aocheng Luo , Dongyue Li , Chang Liu , Sheng Chen , Yangang Zhang , Junzhi Yu

Aerial Vision-and-Language Navigation (VLN) is a novel task enabling Unmanned Aerial Vehicles (UAVs) to navigate in outdoor environments through natural language instructions and visual cues. However, it remains challenging due to the…

Robotics · Computer Science 2025-08-12 Yunpeng Gao , Zhigang Wang , Pengfei Han , Linglin Jing , Dong Wang , Bin Zhao

In recent years, the rapid advancement of Large Language Models (LLMs) such as the Generative Pre-trained Transformer (GPT) has attracted increasing attention due to their potential in a variety of practical applications. The application of…

Artificial Intelligence · Computer Science 2025-04-24 Jinzhou Lin , Han Gao , Xuxiang Feng , Rongtao Xu , Changwei Wang , Man Zhang , Li Guo , Shibiao Xu

Object Goal Navigation (ObjectNav) challenges robots to find objects in unseen environments, demanding sophisticated reasoning. While Vision-Language Models (VLMs) show potential, current ObjectNav methods often employ them superficially,…

Robotics · Computer Science 2025-06-23 Mobin Habibpour , Fatemeh Afghah

Vision-and-Language Navigation (VLN) requires an agent to ground language instructions to its own movement within a visual environment. While state-of-the-art methods leverage the reasoning capabilities of Vision-Language Models (VLMs) for…

We present LGX (Language-guided Exploration), a novel algorithm for Language-Driven Zero-Shot Object Goal Navigation (L-ZSON), where an embodied agent navigates to a uniquely described target object in a previously unseen environment. Our…

Robotics · Computer Science 2024-04-16 Vishnu Sashank Dorbala , James F. Mullen , Dinesh Manocha

Indoor navigation presents unique challenges due to complex layouts and the unavailability of GNSS signals. Existing solutions often struggle with contextual adaptation, and typically require dedicated hardware. In this work, we explore the…

Artificial Intelligence · Computer Science 2025-06-23 Alberto Coffrini , Paolo Barsocchi , Francesco Furfari , Antonino Crivello , Alessio Ferrari

With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achieve open-world visual perception remains an open question. In…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Chris Kelly , Luhui Hu , Bang Yang , Yu Tian , Deshun Yang , Cindy Yang , Zaoshan Huang , Zihao Li , Jiayin Hu , Yuexian Zou

Large language models (LLMs) support data analysis through conversational user interfaces, as exemplified in OpenAI's ChatGPT (formally known as Advanced Data Analysis or Code Interpreter). Essentially, LLMs produce code for accomplishing…

Human-Computer Interaction · Computer Science 2024-08-06 Liwenhan Xie , Chengbo Zheng , Haijun Xia , Huamin Qu , Chen Zhu-Tian

Embodied navigation for long-horizon tasks, guided by complex natural language instructions, remains a formidable challenge in artificial intelligence. Existing agents often struggle with robust long-term planning about unseen environments,…

Robotics · Computer Science 2026-03-16 Fei Liu , Shichao Xie , Minghua Luo , Zedong Chu , Junjun Hu , Xiaolong Wu , Mu Xu

The field of vision-and-language (VL) understanding has made unprecedented progress with end-to-end large pre-trained VL models (VLMs). However, they still fall short in zero-shot reasoning tasks that require multi-step inferencing. To…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Haoxuan You , Zhecan Wang , Rui Sun , Long Chen , Gengyu Wang , Hammad A. Ayyubi , Kai-Wei Chang , Shih-Fu Chang

Visual navigation is an essential skill for home-assistance robots, providing the object-searching ability to accomplish long-horizon daily tasks. Many recent approaches use Large Language Models (LLMs) for commonsense inference to improve…

Robotics · Computer Science 2024-10-15 Xinxin Zhao , Wenzhe Cai , Likun Tang , Teng Wang

Vision-and-language navigation (VLN) stands as a key research problem of Embodied AI, aiming at enabling agents to navigate in unseen environments following linguistic instructions. In this field, generalization is a long-standing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Jiazhao Zhang , Kunyu Wang , Rongtao Xu , Gengze Zhou , Yicong Hong , Xiaomeng Fang , Qi Wu , Zhizheng Zhang , He Wang

Goal-conditioned policies for robotic navigation can be trained on large, unannotated datasets, providing for good generalization to real-world settings. However, particularly in vision-based settings where specifying goals requires an…

Robotics · Computer Science 2022-07-27 Dhruv Shah , Blazej Osinski , Brian Ichter , Sergey Levine

Existing navigation decision support systems often perform poorly when handling non-predefined navigation scenarios. Leveraging the generalization capabilities of large language model (LLM) in handling unknown scenarios, this research…

Artificial Intelligence · Computer Science 2025-02-25 Feng Ma , Xiu-min Wang , Chen Chen , Xiao-bin Xu , Xin-ping Yan

Multimodal large language models (MLLMs) have emerged as a prominent area of interest within the research community, given their proficiency in handling and reasoning with non-textual data, including images and videos. This study seeks to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Zhenhua Xu , Yujia Zhang , Enze Xie , Zhen Zhao , Yong Guo , Kwan-Yee. K. Wong , Zhenguo Li , Hengshuang Zhao

Enabling robotic assistants to navigate complex environments and locate objects described in free-form language is a critical capability for real-world deployment. While foundation models, particularly Vision-Language Models (VLMs), offer…

Robotics · Computer Science 2026-04-16 Naoki Yokoyama , Sehoon Ha