English
Related papers

Related papers: Counterfactual Cycle-Consistent Learning for Instr…

200 papers

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Vision Language Model to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Fan Yang , Zhiyang Chen , Yousong Zhu , Xin Li , Jinqiao Wang

Vision-Language Instruction Tuning (VLIT) is a critical training phase for Large Vision-Language Models (LVLMs). With the improving capabilities of open-source LVLMs, researchers have increasingly turned to generate VLIT data by using…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Ji Ma , Wei Suo , Peng Wang , Yanning Zhang

With the recent success of the pre-training technique for NLP and image-linguistic tasks, some video-linguistic pre-training works are gradually developed to improve video-text related downstream tasks. However, most of the existing…

Computer Vision and Pattern Recognition · Computer Science 2020-09-16 Huaishao Luo , Lei Ji , Botian Shi , Haoyang Huang , Nan Duan , Tianrui Li , Jason Li , Taroon Bharti , Ming Zhou

Natural Language-conditioned reinforcement learning (RL) enables the agents to follow human instructions. Previous approaches generally implemented language-conditioned RL by providing human instructions in natural language (NL) and…

Computation and Language · Computer Science 2023-02-21 Jing-Cheng Pang , Xin-Yu Yang , Si-Hang Yang , Yang Yu

Recent research in Vision Language Navigation (VLN) has overlooked the development of agents' inquisitive abilities, which allow them to ask clarifying questions when instructions are incomplete. This paper addresses how agents can…

Artificial Intelligence · Computer Science 2024-11-12 Savitha Sam Abraham , Sourav Garg , Feras Dayoub

Vision-and-Language Navigation (VLN) is a challenging task where an agent is required to navigate to a natural language described location via vision observations. The navigation abilities of the agent can be enhanced by the relations…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Bowen Huang , Yanwei Zheng , Chuanlin Lan , Xinpeng Zhao , Yifei Zou , Dongxiao yu

To be successful, Vision-and-Language Navigation (VLN) agents must be able to ground instructions to actions based on their surroundings. In this work, we develop a methodology to study agent behavior on a skill-specific basis -- examining…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Zijiao Yang , Arjun Majumdar , Stefan Lee

Vision-Language Navigation requires agents to act coherently over long horizons by understanding not only local visual context but also how far they have advanced within a multi-step instruction. However, recent Vision-Language-Action…

Deep Learning has revolutionized our ability to solve complex problems such as Vision-and-Language Navigation (VLN). This task requires the agent to navigate to a goal purely based on visual sensory inputs given natural language…

Robotics · Computer Science 2021-04-22 Muhammad Zubair Irshad , Chih-Yao Ma , Zsolt Kira

Why must vision-language navigation be bound to detailed and verbose language instructions? While such details ease decision-making, they fundamentally contradict the goal for navigation in the real-world. Ideally, agents should possess the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Hai Zhang , Siqi Liang , Li Chen , Yuxian Li , Yukuan Xu , Yichao Zhong , Fu Zhang , Hongyang Li

Counterfactual explanation is a common class of methods to make local explanations of machine learning decisions. For a given instance, these methods aim to find the smallest modification of feature values that changes the predicted…

Artificial Intelligence · Computer Science 2022-12-22 Victor Guyomard , Françoise Fessant , Thomas Guyet , Tassadit Bouadi , Alexandre Termier

A visually-grounded navigation instruction can be interpreted as a sequence of expected observations and actions an agent following the correct trajectory would encounter and perform. Based on this intuition, we formulate the problem of…

Computer Vision and Pattern Recognition · Computer Science 2019-11-27 Peter Anderson , Ayush Shrivastava , Devi Parikh , Dhruv Batra , Stefan Lee

Data scarcity is a long-standing challenge in the Vision-Language Navigation (VLN) field, which extremely hinders the generalization of agents to unseen environments. Previous works primarily rely on additional simulator data or…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Ziming Wei , Bingqian Lin , Yunshuang Nie , Jiaqi Chen , Shikui Ma , Hang Xu , Xiaodan Liang

Natural language understanding (NLU) and natural language generation (NLG) are two fundamental and related tasks in building task-oriented dialogue systems with opposite objectives: NLU tackles the transformation from natural language to…

Computation and Language · Computer Science 2020-06-16 Bo-Hsiang Tseng , Jianpeng Cheng , Yimai Fang , David Vandyke

Understanding and following directions provided by humans can enable robots to navigate effectively in unknown situations. We present FollowNet, an end-to-end differentiable neural architecture for learning multi-modal navigation policies.…

Robotics · Computer Science 2018-09-20 Pararth Shah , Marek Fiser , Aleksandra Faust , J. Chase Kew , Dilek Hakkani-Tur

Recently, Vision-Language-Action (VLA) models have demonstrated strong performance on a range of robotic tasks. These models rely on multimodal inputs, with language instructions playing a crucial role -- not only in predicting actions, but…

Artificial Intelligence · Computer Science 2025-08-25 Wen-Han Hsieh , Elvis Hsieh , Dantong Niu , Trevor Darrell , Roei Herzig , David M. Chan

The aspiration of the Vision-and-Language Navigation (VLN) task has long been to develop an embodied agent with robust adaptability, capable of seamlessly transferring its navigation capabilities across various tasks. Despite remarkable…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Siqi Zhang , Yanyuan Qiao , Qunbo Wang , Longteng Guo , Zhihua Wei , Jing Liu

Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing methods, whether rule-based or data-driven, frequently struggle to capture complex scene…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Kefei Tian , Yuansheng Lian , Kai Yang , Xiangdong Chen , Shen Li

Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about actions and temporal order. Existing mitigation strategies, such as textual filtering or random video…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Tobia Poppi , Burak Uzkent , Amanmeet Garg , Lucas Porto , Garin Kessler , Yezhou Yang , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara , Florian Schiffers

Vision-Language Navigation requires the agent to follow natural language instructions to reach a specific target. The large discrepancy between seen and unseen environments makes it challenging for the agent to generalize well. Previous…

Computer Vision and Pattern Recognition · Computer Science 2022-09-13 Yujie Lu , Huiliang Zhang , Ping Nie , Weixi Feng , Wenda Xu , Xin Eric Wang , William Yang Wang