English
Related papers

Related papers: Think, Remember, Navigate: Zero-Shot Object-Goal N…

200 papers

Vision-and-Language Navigation (VLN) poses significant challenges for agents to interpret natural language instructions and navigate complex 3D environments. While recent progress has been driven by large-scale pre-training and data…

Artificial Intelligence · Computer Science 2026-05-14 Tianyi Ma , Yue Zhang , Zehao Wang , Parisa Kordjamshidi

Navigating to a specified object in an unknown environment is a fundamental yet challenging capability of embodied intelligence. However, current methods struggle to balance decision frequency with intelligence, resulting in decisions…

VLA models have shown promising potential in embodied navigation by unifying perception and planning while inheriting the strong generalization abilities of large VLMs. However, most existing VLA models rely on reactive mappings directly…

Robotics · Computer Science 2026-01-14 Shaoan Wang , Yuanfei Luo , Xingyu Chen , Aocheng Luo , Dongyue Li , Chang Liu , Sheng Chen , Yangang Zhang , Junzhi Yu

Vision-language models (VLMs) have emerged as a promising direction for end-to-end autonomous driving (AD) by jointly modeling visual observations, driving context, and language-based reasoning. However, existing VLM-based systems face a…

Robotics · Computer Science 2026-03-10 Ximeng Tao , Pardis Taghavi , Dimitar Filev , Reza Langari , Gaurav Pandey

In this article, we investigate vision-language models (VLM) as reasoners. The ability to form abstractions underlies mathematical reasoning, problem-solving, and other Math AI tasks. Several formalisms have been given to these underlying…

Artificial Intelligence · Computer Science 2024-07-08 Denisa Roberts , Lucas Roberts

Vision-and-Language Navigation (VLN), as a crucial research problem of Embodied AI, requires an embodied agent to navigate through complex 3D environments following natural language instructions. Recent research has highlighted the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Bingqian Lin , Yunshuang Nie , Ziming Wei , Jiaqi Chen , Shikui Ma , Jianhua Han , Hang Xu , Xiaojun Chang , Xiaodan Liang

The deployment of artificial intelligence models at the edge is increasingly critical for autonomous robots operating in GPS-denied environments where local, resource-efficient reasoning is essential. This work demonstrates the feasibility…

Robotics · Computer Science 2025-11-11 Justin Williams , Kishor Datta Gupta , Roy George , Mrinmoy Sarkar

Recent research looks to harness the general knowledge and reasoning of large language models (LLMs) into agents that accomplish user-specified goals in interactive environments. Vision-language models (VLMs) extend LLMs to multi-modal data…

Machine Learning · Computer Science 2025-05-07 Jake Grigsby , Yuke Zhu , Michael Ryoo , Juan Carlos Niebles

Vision-and-Language Navigation (VLN) requires an agent to dynamically explore complex 3D environments following human instructions. Recent research underscores the potential of harnessing large language models (LLMs) for VLN, given their…

Artificial Intelligence · Computer Science 2025-11-19 Yu Zhong , Zihao Zhang , Rui Zhang , Lingdong Huang , Haihan Gao , Shuo Wang , Da Li , Ruijian Han , Jiaming Guo , Shaohui Peng , Di Huang , Yunji Chen

Vision-language navigation (VLN) is a critical domain within embedded intelligence, requiring agents to navigate 3D environments based on natural language instructions. Traditional VLN research has focused on improving environmental…

Artificial Intelligence · Computer Science 2024-09-24 Zhiyuan Li , Yanfeng Lv , Ziqin Tu , Di Shang , Hong Qiao

Conventional approaches to vision-and-language navigation (VLN) are trained end-to-end but struggle to perform well in freely traversable environments. Inspired by the robotics community, we propose a modular approach to VLN using…

Robotics · Computer Science 2020-12-11 Kevin Chen , Junshen K. Chen , Jo Chuang , Marynel Vázquez , Silvio Savarese

Visual Language Models (VLMs) are essential for various tasks, particularly visual reasoning tasks, due to their robust multi-modal information integration, visual reasoning capabilities, and contextual awareness. However, existing \VLMs{}'…

Computation and Language · Computer Science 2024-09-13 Zaiqiao Meng , Hao Zhou , Yifang Chen

Visual reasoning is dominated by end-to-end neural networks scaled to billions of model parameters and training examples. However, even the largest models struggle with compositional reasoning, generalization, fine-grained spatial and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Aleksandar Stanić , Sergi Caelles , Michael Tschannen

Bridging the gap between embodied intelligence and embedded deployment remains a key challenge in intelligent robotic systems, where perception, reasoning, and planning must operate under strict constraints on computation, memory, energy,…

Robotics · Computer Science 2026-05-19 Kuan Xu , Ruimeng Liu , Yizhuo Yang , Denan Liang , Tongxing Jin , Shenghai Yuan , Chen Wang , Lihua Xie

When navigating in a man-made environment they haven't visited before--like an office building--humans employ behaviors such as reading signs and asking others for directions. These behaviors help humans reach their destinations efficiently…

Robotics · Computer Science 2025-09-26 Bhargav Chandaka , Gloria X. Wang , Haozhe Chen , Henry Che , Albert J. Zhai , Shenlong Wang

Pre-trained large language models (LLMs) have demonstrated strong common-sense reasoning abilities, making them promising for robotic navigation and planning tasks. However, despite recent progress, bridging the gap between language…

Robotics · Computer Science 2025-12-29 Mingfeng Yuan , Letian Wang , Steven L. Waslander

Language-guided embodied navigation requires an agent to interpret object-referential instructions, search across multiple rooms, localize the referenced target, and execute reliable motion toward it. Existing systems remain limited in real…

Robotics · Computer Science 2026-03-19 Zhongyuang Liu , Min He , Shaonan Yu , Xinhang Xu , Muqing Cao , Jianping Li , Jianfei Yang , Lihua Xie

Current Visual-Language Navigation (VLN) methodologies face a trade-off between semantic understanding and control precision. While Multimodal Large Language Models (MLLMs) offer superior reasoning, deploying them as low-level controllers…

Robotics · Computer Science 2026-02-19 Zhenxing Xu , Brikit Lu , Weidong Bao , Zhengqiu Zhu , Junsong Zhang , Hui Yan , Wenhao Lu , Ji Wang

Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two challenges: a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Zuntao Liu , Yi Du , Taimeng Fu , Shaoshu Su , Cherie Ho , Chen Wang

Vision-Language Models (VLMs) have recently gained attention due to their competitive performance on multiple downstream tasks, achieved by following user-input instructions. However, VLMs still exhibit several limitations in visual…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Simone Alghisi , Gabriel Roccabruna , Massimo Rizzoli , Seyed Mahed Mousavi , Giuseppe Riccardi