English
Related papers

Related papers: Generative Language-Grounded Policy in Vision-and-…

200 papers

The existing methods for Vision and Language Navigation in the Continuous Environment (VLN-CE) commonly incorporate a waypoint predictor to discretize the environment. This simplifies the navigation actions into a view selection task and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Yue Zhang , Parisa Kordjamshidi

Predicting temporal progress from visual trajectories is important for intelligent robots that can learn, adapt, and improve. However, learning such progress estimator, or temporal value function, across different tasks and domains requires…

Vision-language models (VLMs) have tremendous potential for grounding language, and thus enabling language-conditioned agents (LCAs) to perform diverse tasks specified with text. This has motivated the study of LCAs based on reinforcement…

Artificial Intelligence · Computer Science 2024-11-27 Theo Cachet , Christopher R. Dance , Olivier Sigaud

Vision-Language-Action (VLA) models are emerging as highly effective planning models for end-to-end autonomous driving systems. However, current works mostly rely on imitation learning from sparse trajectory annotations and under-utilize…

This paper explores the possibility of learning custom tokens for representing new concepts in Vision-Language Models (VLMs). Our aim is to learn tokens that can be effective for both discriminative and generative tasks while composing well…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Pramuditha Perera , Matthew Trager , Luca Zancato , Alessandro Achille , Stefano Soatto

Vision-and-Language Navigation (VLN) is an essential skill for embodied agents, allowing them to navigate in 3D environments following natural language instructions. High-performance navigation models require a large amount of training…

Artificial Intelligence · Computer Science 2025-03-10 Zihan Wang , Yaohui Zhu , Gim Hee Lee , Yachun Fan

The Object Goal Navigation (ObjectNav) task challenges agents to locate a specified object in an unseen environment by imagining unobserved regions of the scene. Prior approaches rely on deterministic and discriminative models to complete…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Badi Li , Ren-jie Lu , Yu Zhou , Jingke Meng , Wei-shi Zheng

Vision-language models (VLMs) have been widely-applied in ground-based vision-language navigation (VLN). However, the vast complexity of outdoor aerial environments compounds data acquisition challenges and imposes long-horizon trajectory…

Robotics · Computer Science 2025-11-24 Peican Lin , Gan Sun , Chenxi Liu , Fazeng Li , Weihong Ren , Yang Cong

Accuracy of many visiolinguistic tasks has benefited significantly from the application of vision-and-language(V&L) BERT. However, its application for the task of vision-and-language navigation (VLN) remains limited. One reason for this is…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Yicong Hong , Qi Wu , Yuankai Qi , Cristian Rodriguez-Opazo , Stephen Gould

We present Vision-based Navigation with Language-based Assistance (VNLA), a grounded vision-language task where an agent with visual perception is guided via language to find objects in photorealistic indoor environments. The task emulates…

Machine Learning · Computer Science 2019-04-09 Khanh Nguyen , Debadeepta Dey , Chris Brockett , Bill Dolan

Most existing works solving Room-to-Room VLN problem only utilize RGB images and do not consider local context around candidate views, which lack sufficient visual cues about surrounding environment. Moreover, natural language contains…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Jingyang Huo , Qiang Sun , Boyan Jiang , Haitao Lin , Yanwei Fu

Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, generalization remains a persistent challenge, particularly when…

Robotics · Computer Science 2025-02-27 Zerui Li , Gengze Zhou , Haodong Hong , Yanyan Shao , Wenqi Lyu , Yanyuan Qiao , Qi Wu

The study of vision-and-language navigation (VLN) has typically relied on expert trajectories, which may not always be available in real-world situations due to the significant effort required to collect them. On the other hand, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Valay Bundele , Mahesh Bhupati , Biplab Banerjee , Aditya Grover

Vision-language models (VLMs) have recently emerged as powerful representation learning systems that align visual observations with natural language concepts, offering new opportunities for semantic reasoning in safety-critical autonomous…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Ross Greer , Maitrayee Keskar , Angel Martinez-Sanchez , Parthib Roy , Shashank Shriram , Mohan Trivedi

Vision and Language Navigation (VLN) is a challenging task that requires agents to understand instructions and navigate to the destination in a visual environment.One of the key challenges in outdoor VLN is keeping track of which part of…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Huilin Tian , Jingke Meng , Wei-Shi Zheng , Yuan-Ming Li , Junkai Yan , Yunong Zhang

Evaluating the surroundings to gain understanding, frame perspectives, and anticipate behavioral reactions is an inherent human trait. However, these continuous encounters are diverse and complex, posing challenges to their study and…

Computers and Society · Computer Science 2026-02-26 Deepank Verma , Olaf Mumm , Vanessa Miriam Carlow

Vision-and-Language Navigation (VLN) has long been constrained by the limited diversity and scalability of simulator-curated datasets, which fail to capture the complexity of real-world environments. To overcome this limitation, we…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Mingfei Han , Haihong Hao , Liang Ma , Kamila Zhumakhanova , Ekaterina Radionova , Jingyi Zhang , Xiaojun Chang , Xiaodan Liang , Ivan Laptev

Real-world sequential decision making is characterized by sparse rewards and large decision spaces, posing significant difficulty for experiential learning systems like $\textit{tabula rasa}$ reinforcement learning (RL) agents. Large…

Computation and Language · Computer Science 2024-03-06 Hitesh Golchha , Sahil Yerawar , Dhruvesh Patel , Soham Dan , Keerthiram Murugesan

Integrating large language models (LLMs) into embodied AI models is becoming increasingly prevalent. However, existing zero-shot LLM-based Vision-and-Language Navigation (VLN) agents either encode images as textual scene descriptions,…

Artificial Intelligence · Computer Science 2025-09-30 Yue Zhang , Tianyi Ma , Zun Wang , Yanyuan Qiao , Parisa Kordjamshidi

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location in 3D environments following the natural language instruction. In this field, the agent is usually trained and evaluated in the navigation simulators,…

Robotics · Computer Science 2024-10-15 Zihan Wang , Xiangyang Li , Jiahao Yang , Yeqi Liu , Shuqiang Jiang