English
Related papers

Related papers: World-Consistent Data Generation for Vision-and-La…

200 papers

Vision-Language-Navigation (VLN) models exhibit excellent navigation accuracy but incur high computational overhead. Token caching has emerged as a promising training-free strategy to reduce this cost by reusing token computation results;…

Vision-and-language navigation (VLN) is a long-standing challenge in autonomous robotics, aiming to empower agents with the ability to follow human instructions while navigating complex environments. Two key bottlenecks remain in this…

Robotics · Computer Science 2025-06-13 Yuhang Zhang , Haosheng Yu , Jiaping Xiao , Mir Feroskhan

Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has propelled advancements in VLN, the significance of foreground…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Yunbo Xu , Xuesong Zhang , Jia Li , Zhenzhen Hu , Richang Hong

Core to the vision-and-language navigation (VLN) challenge is building robust instruction representations and action decoding schemes, which can generalize well to previously unseen instructions and environments. In this paper, we report…

Computation and Language · Computer Science 2019-09-06 Xiujun Li , Chunyuan Li , Qiaolin Xia , Yonatan Bisk , Asli Celikyilmaz , Jianfeng Gao , Noah Smith , Yejin Choi

Embodied navigation for long-horizon tasks, guided by complex natural language instructions, remains a formidable challenge in artificial intelligence. Existing agents often struggle with robust long-term planning about unseen environments,…

Robotics · Computer Science 2026-03-16 Fei Liu , Shichao Xie , Minghua Luo , Zedong Chu , Junjun Hu , Xiaolong Wu , Mu Xu

Reliable anticipation of traffic accidents is essential for advancing autonomous driving systems. However, this objective is limited by two fundamental challenges: the scarcity of diverse, high-quality training data and the frequent absence…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Yanchen Guan , Haicheng Liao , Chengyue Wang , Xingcheng Liu , Jiaxun Zhang , Zhenning Li

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse driving videos, while preserving the spatiotemporal…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Li-Heng Chen , Ke Cheng , Yahui Liu , Lei Shi , Shi-Sheng Huang , Hongbo Fu

Recently, numerous algorithms have been developed to tackle the problem of vision-language navigation (VLN), i.e., entailing an agent to navigate 3D environments through following linguistic instructions. However, current VLN agents simply…

Computer Vision and Pattern Recognition · Computer Science 2021-03-08 Hanqing Wang , Wenguan Wang , Wei Liang , Caiming Xiong , Jianbing Shen

Vision-Language Models (VLMs) often yield inconsistent descriptions of the same object across viewpoints, hindering the ability of embodied agents to construct consistent semantic representations over time. Previous methods resolved…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Tommaso Galliena , Stefano Rosa , Tommaso Apicella , Pietro Morerio , Alessio Del Bue , Lorenzo Natale

Vision-Language Navigation (VLN) requires the agent to follow language instructions to reach a target position. A key factor for successful navigation is to align the landmarks implied in the instruction with diverse visual observations.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Bingqian Lin , Yunshuang Nie , Ziming Wei , Yi Zhu , Hang Xu , Shikui Ma , Jianzhuang Liu , Xiaodan Liang

Vision-and-language navigation (VLN) agents are trained to navigate in real-world environments by following natural language instructions. A major challenge in VLN is the limited availability of training data, which hinders the models'…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Zi-Yi Dou , Feng Gao , Nanyun Peng

In the Vision-and-Language Navigation (VLN) task, an agent with egocentric vision navigates to a destination given natural language instructions. The act of manually annotating these instructions is timely and expensive, such that many…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Felix Yu , Zhiwei Deng , Karthik Narasimhan , Olga Russakovsky

Developing Vision-and-Language Navigation (VLN) agents typically assumes a \textit{train-once-deploy-once} strategy, which is unrealistic as deployed agents continually encounter novel environments. To address this, we propose the Continual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Seongjun Jeong , Gi-Cheon Kang , Seongho Choi , Joochan Kim , Byoung-Tak Zhang

Unifying diverse image generation tasks within a single framework remains a fundamental challenge in visual generation. While large language models (LLMs) achieve unification through task-agnostic data and generation, existing visual…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yijing Lin , Mengqi Huang , Shuhan Zhuang , Zhendong Mao

Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal to Noise Ratio (PSNR) reward pixel fidelity over semantic correctness, while Frechet…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Karthik Inbasekar , Guy Rom , Omer Shlomovits

Autonomous driving systems have made significant advances in Q&A, perception, prediction, and planning based on local visual information, yet they struggle to incorporate broader navigational context that human drivers routinely utilize. We…

Robotics · Computer Science 2025-11-04 Qucheng Peng , Chen Bai , Guoxiang Zhang , Bo Xu , Xiaotong Liu , Xiaoyin Zheng , Chen Chen , Cheng Lu

Vision-language models (VLMs) are highly effective but often underperform on specialized tasks; for example, Llava-1.5 struggles with chart and diagram understanding due to scarce task-specific training data. Existing training data, sourced…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Siddharth Joshi , Besmira Nushi , Vidhisha Balachandran , Varun Chandrasekaran , Vibhav Vineet , Neel Joshi , Baharan Mirzasoleiman

Vision-and-Language Navigation (VLN) requires agents to interpret natural language instructions and act coherently in visually rich environments. However, most existing methods rely on reactive state-action mappings without explicitly…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Weiye Zhu , Zekai Zhang , Xiangchen Wang , Hewei Pan , Teng Wang , Tiantian Geng , Rongtao Xu , Feng Zheng

Most existing works in vision-and-language navigation (VLN) focus on either discrete or continuous environments, training agents that cannot generalize across the two. The fundamental difference between the two setups is that discrete…

Computer Vision and Pattern Recognition · Computer Science 2022-03-08 Yicong Hong , Zun Wang , Qi Wu , Stephen Gould

Why must vision-language navigation be bound to detailed and verbose language instructions? While such details ease decision-making, they fundamentally contradict the goal for navigation in the real-world. Ideally, agents should possess the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Hai Zhang , Siqi Liang , Li Chen , Yuxian Li , Yukuan Xu , Yichao Zhong , Fu Zhang , Hongyang Li