English
Related papers

Related papers: SeqWalker: Sequential-Horizon Vision-and-Language …

200 papers

Vision and Language Navigation (VLN) requires an agent to navigate to a target location by following natural language instructions. Most of existing works represent a navigation candidate by the feature of the corresponding single view…

Computer Vision and Pattern Recognition · Computer Science 2021-07-27 Dong An , Yuankai Qi , Yan Huang , Qi Wu , Liang Wang , Tieniu Tan

Vision-language models (VLMs) have demonstrated remarkable potential in integrating visual and linguistic information, but their performance is often constrained by the need for extensive, high-quality image-text training data. Curation of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Giorgio Giannone , Ruoteng Li , Qianli Feng , Evgeny Perevodchikov , Rui Chen , Aleix Martinez

Vision-and-Language Navigation (VLN) is a task that an agent is required to follow a language instruction to navigate to the goal position, which relies on the ongoing interactions with the environment during moving. Recent…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Chuang Lin , Yi Jiang , Jianfei Cai , Lizhen Qu , Gholamreza Haffari , Zehuan Yuan

Large language models (LLMs) have shown remarkable advancements in enabling language agents to tackle simple tasks. However, applying them for complex, multi-step, long-horizon tasks remains a challenge. Recent work have found success by…

Computation and Language · Computer Science 2025-08-05 Lutfi Eren Erdogan , Nicholas Lee , Sehoon Kim , Suhong Moon , Hiroki Furuta , Gopala Anumanchipalli , Kurt Keutzer , Amir Gholami

Multimodal large language models (MLLMs) have shown promising potential in Vision-Language Navigation (VLN). However, their practical development is severely hindered by the substantial training overhead. We recognize two key issues that…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Duo Zheng , Shijia Huang , Yanyang Li , Liwei Wang

Since the rise of vision-language navigation (VLN), great progress has been made in instruction following -- building a follower to navigate environments under the guidance of instructions. However, far less attention has been paid to the…

Computer Vision and Pattern Recognition · Computer Science 2022-04-01 Hanqing Wang , Wei Liang , Jianbing Shen , Luc Van Gool , Wenguan Wang

Vision Language Navigation (VLN) typically requires agents to navigate to specified objects or remote regions in unknown scenes by obeying linguistic commands. Such tasks require organizing historical visual observations for linguistic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Bolei Chen , Jiaxu Kang , Yifei Wang , Ping Zhong , Qi Wu , Jianxin Wang

In this work, we propose a modular approach for the Vision-Language Navigation (VLN) task by decomposing the problem into four sub-modules that use state-of-the-art Large Language Models (LLMs) and Vision-Language Models (VLMs) in a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Navid Rajabi , Jana Kosecka

Vision-language pre-training (VLP) on large-scale image-text pairs has recently witnessed rapid progress for learning cross-modal representations. Existing pre-training methods either directly concatenate image representation and text…

Computation and Language · Computer Science 2021-03-16 Chenliang Li , Ming Yan , Haiyang Xu , Fuli Luo , Wei Wang , Bin Bi , Songfang Huang

Vision-and-Language Navigation (VLN) requires an embodied agent to traverse complex environments by following natural language instructions, demanding accurate alignment between visual observations and linguistic guidance. Despite recent…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Yaohua Liu , Xinyuan Song , Yunfu Deng , Yifan Xie , Binkai Ou , Yan Zhong

The increasingly complex and diverse planetary exploration environment requires more adaptable and flexible rover navigation strategy. In this study, we propose a VLM-empowered multi-mode system to achieve efficient while safe autonomous…

Robotics · Computer Science 2025-06-23 Sinuo Cheng , Ruyi Zhou , Wenhao Feng , Huaiguang Yang , Haibo Gao , Zongquan Deng , Liang Ding

Despite recent advancements in Multi-modal Large Language Models (MLLMs) on diverse understanding tasks, these models struggle to solve problems which require extensive multi-step reasoning. This is primarily due to the progressive dilution…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Byungwoo Jeon , Yoonwoo Jeong , Hyunseok Lee , Minsu Cho , Jinwoo Shin

Conventional Vision-and-Language Navigation (VLN) benchmarks assume instructions are feasible and the referenced target exists, leaving agents ill-equipped to handle false-premise goals. We introduce VLN-NF, a benchmark with false-premise…

Robotics · Computer Science 2026-04-21 Hung-Ting Su , Ting-Jun Wang , Jia-Fong Yeh , Min Sun , Winston H. Hsu

Vision-and-language navigation (VLN) simulates a visual agent that follows natural-language navigation instructions in real-world scenes. Existing approaches have made enormous progress in navigation in new environments, such as beam…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Qi Zheng , Daqing Liu , Chaoyue Wang , Jing Zhang , Dadong Wang , Dacheng Tao

Vision-language navigation is a task that requires an agent to follow instructions to navigate in environments. It becomes increasingly crucial in the field of embodied AI, with potential applications in autonomous navigation, search and…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Dong An , Hanqing Wang , Wenguan Wang , Zun Wang , Yan Huang , Keji He , Liang Wang

Autonomous navigation in partially observable environments requires agents to reason beyond immediate sensor input, exploit occlusion, and ensure safety while progressing toward a goal. These challenges arise in many robotics domains, from…

Robotics · Computer Science 2026-04-21 Mihir Chauhan , Damon Conover , Aniket Bera

Language-guided embodied navigation requires an agent to interpret object-referential instructions, search across multiple rooms, localize the referenced target, and execute reliable motion toward it. Existing systems remain limited in real…

Robotics · Computer Science 2026-03-19 Zhongyuang Liu , Min He , Shaonan Yu , Xinhang Xu , Muqing Cao , Jianping Li , Jianfei Yang , Lihua Xie

Existing UAV vision-language navigation (VLN) benchmarks have enabled language-guided flight, but they largely focus on long, step-wise route descriptions with goal-centric evaluation, making them less diagnostic for real operations where…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Jingyu Guo , Ziye Chen , Ziwen Li , Zhengqing Gao , Jiaxin Huang , Hanlue Zhang , Fengming Huang , Yu Yao , Tongliang Liu , Mingming Gong

Learning to navigate in a visual environment following natural-language instructions is a challenging task, because the multimodal inputs to the agent are highly variable, and the training data on a new task is often limited. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2020-04-07 Weituo Hao , Chunyuan Li , Xiujun Li , Lawrence Carin , Jianfeng Gao

In most existing embodied navigation tasks, instructions are well-defined and unambiguous, such as instruction following and object searching. Under this idealized setting, agents are required solely to produce effective navigation outputs…

Robotics · Computer Science 2026-01-26 Wensi Huang , Shaohao Zhu , Meng Wei , Jinming Xu , Xihui Liu , Hanqing Wang , Tai Wang , Feng Zhao , Jiangmiao Pang