English
Related papers

Related papers: End-to-End Navigation with Vision Language Models:…

200 papers

Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture temporal dynamics and physical laws. We propose NavDreamer,…

Robotics · Computer Science 2026-02-11 Xijie Huang , Weiqi Gai , Tianyue Wu , Congyu Wang , Zhiyang Liu , Xin Zhou , Yuze Wu , Fei Gao

Exploring open-world situations in an end-to-end manner is a promising yet challenging task due to the need for strong generalization capabilities. In particular, end-to-end autonomous driving in unstructured outdoor environments often…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Hyunki Seong , Seongwoo Moon , Hojin Ahn , Jehun Kang , David Hyunchul Shim

Flowcharts are typically presented as images, driving the trend of using vision-language models (VLMs) for end-to-end flowchart understanding. However, two key challenges arise: (i) Limited controllability--users have minimal influence over…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Junyi Ye , Ankan Dash , Wenpeng Yin , Guiling Wang

Vision-language models (VLMs) have been widely-applied in ground-based vision-language navigation (VLN). However, the vast complexity of outdoor aerial environments compounds data acquisition challenges and imposes long-horizon trajectory…

Robotics · Computer Science 2025-11-24 Peican Lin , Gan Sun , Chenxi Liu , Fazeng Li , Weihong Ren , Yang Cong

Urban monitoring of public infrastructure (such as waste bins, road signs, vegetation, sidewalks, and construction sites) poses significant challenges due to the diversity of objects, environments, and contextual conditions involved.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 André Torneiro , Diogo Monteiro , Paulo Novais , Pedro Rangel Henriques , Nuno F. Rodrigues

Vision-Language Models (VLMs) have been increasingly integrated into object navigation tasks for their rich prior knowledge and strong reasoning abilities. However, applying VLMs to navigation poses two key challenges: effectively…

Robotics · Computer Science 2025-09-17 Haokun Zhu , Zongtai Li , Zhixuan Liu , Wenshan Wang , Ji Zhang , Jonathan Francis , Jean Oh

Aerial vision-language navigation (AVLN) enables UAVs to follow natural-language instructions in complex 3D environments. However, existing zero-shot AVLN methods often suffer from unstable single-stream Vision-Language Model…

Robotics · Computer Science 2026-03-12 Guiyong Zheng , Yueting Ban , Mingjie Zhang , Juepeng Zheng , Boyu Zhou

Vision-and-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) represents a pivotal challenge in embodied artificial intelligence, focused on enabling UAVs to interpret high-level human commands and execute long-horizon tasks in…

Vision-and-Language Navigation (VLN) is shifting from rigid, step-by-step instruction following toward open-vocabulary, goal-oriented autonomy. Achieving this transition without exhaustive routing prompts requires agents to leverage…

Robotics · Computer Science 2026-03-20 Zihui Yu , Pingcong Li , Bichi Zhang , Sören Schwertfeger

Vehicle-to-everything (V2X) cooperation has emerged as a promising paradigm to overcome the perception limitations of classical autonomous driving by leveraging information from both ego-vehicle and infrastructure sensors. However,…

Robotics · Computer Science 2025-06-23 Junwei You , Haotian Shi , Zhuoyu Jiang , Zilin Huang , Rui Gan , Keshu Wu , Xi Cheng , Xiaopeng Li , Bin Ran

Existing Vision-Language Navigation (VLN) task requires agents to follow verbose instructions, ignoring some potentially useful global spatial priors, limiting their capability to reason about spatial structures. Although human-readable…

Robotics · Computer Science 2026-03-19 Kehan Chen , Yan Huang , Dong An , Jiawei He , Yifei Su , Jing Liu , Nianfeng Liu , Liang Wang

The dominant paradigm for training Large Vision-Language Models (LVLMs) in navigation relies on imitating expert trajectories. This approach reduces the complex navigation task to a sequence-to-sequence replication of a single correct path,…

Robotics · Computer Science 2026-03-24 LinFeng Li , Jian Zhao , Yuan Xie , Xin Tan , Xuelong Li

Recent advances in large Vision-Language Models (VLMs) have provided rich semantic understanding that empowers drones to search for open-set objects via natural language instructions. However, prior systems struggle to integrate VLMs into…

Robotics · Computer Science 2026-01-21 Xuecheng Chen , Zongzhuo Liu , Jianfa Ma , Bang Du , Tiantian Zhang , Xueqian Wang , Boyu Zhou

Unmanned Aerial Vehicle (UAV) Vision-and-Language Navigation (VLN) is vital for applications such as disaster response, logistics delivery, and urban inspection. However, existing methods often struggle with insufficient multimodal fusion,…

Computation and Language · Computer Science 2025-05-20 Hengxing Cai , Jinhan Dong , Jingjun Tan , Jingcheng Deng , Sihang Li , Zhifeng Gao , Haidong Wang , Zicheng Su , Agachai Sumalee , Renxin Zhong

Understanding and following natural language instructions while navigating through complex, real-world environments poses a significant challenge for general-purpose robots. These environments often include obstacles and pedestrians, making…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Xiwen Liang , Liang Ma , Shanshan Guo , Jianhua Han , Hang Xu , Shikui Ma , Xiaodan Liang

While Visual Large Language Models (VLLMs) show great promise as embodied agents, they continue to face substantial challenges in spatial reasoning. Existing embodied benchmarks largely focus on passive, static household environments and…

Robotics · Computer Science 2025-11-24 Yifan Li , Lichi Li , Anh Dao , Xinyu Zhou , Yicheng Qiao , Zheda Mai , Daeun Lee , Zichen Chen , Zhen Tan , Mohit Bansal , Yu Kong

Vision-and-language navigation (VLN) stands as a key research problem of Embodied AI, aiming at enabling agents to navigate in unseen environments following linguistic instructions. In this field, generalization is a long-standing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Jiazhao Zhang , Kunyu Wang , Rongtao Xu , Gengze Zhou , Yicong Hong , Xiaomeng Fang , Qi Wu , Zhizheng Zhang , He Wang

Vision-Language Models (VLMs) offer a promising approach to end-to-end autonomous driving due to their human-like reasoning capabilities. However, troublesome gaps remains between current VLMs and real-world autonomous driving applications.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Hao Jiang , Chuan Hu , Yukang Shi , Yuan He , Ke Wang , Xi Zhang , Zhipeng Zhang

Recent advances in vision-language navigation (VLN) were mainly attributed to emerging large language models (LLMs). These methods exhibited excellent generalization capabilities in instruction understanding and task reasoning. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Xiangyu Dong , Haoran Zhao , Jiang Gao , Haozhou Li , Xiaoguang Ma , Yaoming Zhou , Fuhai Chen , Juan Liu

We introduce VLM-Lens, a toolkit designed to enable systematic benchmarking, analysis, and interpretation of vision-language models (VLMs) by supporting the extraction of intermediate outputs from any layer during the forward pass of…

Computation and Language · Computer Science 2025-10-03 Hala Sheta , Eric Huang , Shuyu Wu , Ilia Alenabi , Jiajun Hong , Ryker Lin , Ruoxi Ning , Daniel Wei , Jialin Yang , Jiawei Zhou , Ziqiao Ma , Freda Shi
‹ Prev 1 8 9 10 Next ›