English
Related papers

Related papers: Planning with the Views via Scene Self-Exploration

200 papers

In this work we propose a holistic framework for autonomous aerial inspection tasks, using semantically-aware, yet, computationally efficient planning and mapping algorithms. The system leverages state-of-the-art receding horizon…

Existing vision-and-language navigation (VLN) models primarily reason over past and current visual observations, while largely ignoring the future visual dynamics induced by actions. As a result, they often lack an effective understanding…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Haihong Hao , Lei Chen , Mingfei Han , Changlin Li , Dong An , Yuqiang Yang , Zhihui Li , Xiaojun Chang

The advent of Vision Language Models (VLM) has allowed researchers to investigate the visual understanding of a neural network using natural language. Beyond object classification and detection, VLMs are capable of visual comprehension and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Haz Sameen Shahgir , Khondker Salman Sayeed , Abhik Bhattacharjee , Wasi Uddin Ahmad , Yue Dong , Rifat Shahriyar

Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Zhihao Yuan , Yibo Peng , Jinke Ren , Yinghong Liao , Yatong Han , Chun-Mei Feng , Hengshuang Zhao , Guanbin Li , Shuguang Cui , Zhen Li

To autonomously navigate and plan interactions in real-world environments, robots require the ability to robustly perceive and map complex, unstructured surrounding scenes. Besides building an internal representation of the observed scene…

The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic understanding abilities, which are essential for handling complex decision-making and long-tail…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Thomas Monninger , Shaoyuan Xie , Qi Alfred Chen , Sihao Ding

Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following human instruction. Compared to ground-based VLN, aerial VLN requires the agent to decide the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Ganlong Zhao , Guanbin Li , Jia Pan , Yizhou Yu

Next Location Prediction is a fundamental task in the study of human mobility, with wide-ranging applications in transportation planning, urban governance, and epidemic forecasting. In practice, when humans attempt to predict the next…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Ruixing Zhang , Yang Zhang , Tongyu Zhu , Leilei Sun , Weifeng Lv

Predicting how the world can evolve in the future is crucial for motion planning in autonomous systems. Classical methods are limited because they rely on costly human annotations in the form of semantic class labels, bounding boxes, and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Tarasha Khurana , Peiyun Hu , David Held , Deva Ramanan

This paper introduces VisionPAD, a novel self-supervised pre-training paradigm designed for vision-centric algorithms in autonomous driving. In contrast to previous approaches that employ neural rendering with explicit depth supervision,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Haiming Zhang , Wending Zhou , Yiyao Zhu , Xu Yan , Jiantao Gao , Dongfeng Bai , Yingjie Cai , Bingbing Liu , Shuguang Cui , Zhen Li

Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception. However, most existing…

Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose the Dynamic 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zihan Wang , Seungjun Lee , Guangzhao Dai , Gim Hee Lee

An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relative to an object B encountered later. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Sahithya Ravi , Gabriel Sarch , Vibhav Vineet , Andrew D. Wilson , Balasaravanan Thoravi Kumaravel

Large vision-language models (VLMs) for autonomous driving (AD) are evolving beyond perception and cognition tasks toward motion planning. However, we identify two critical challenges in this direction: (1) VLMs tend to learn shortcuts by…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Yue Li , Meng Tian , Dechang Zhu , Jiangtong Zhu , Zhenyu Lin , Zhiwei Xiong , Xinhai Zhao

Autonomous exploration of unknown environments is a key capability for mobile robots, but it is largely unsolved for robots equipped with only a single monocular camera and no dense range sensors. In this paper, we present a novel approach…

Robotics · Computer Science 2026-05-15 Tomáš Musil , Matěj Petrlík , Martin Saska

Robotic manipulation requires sophisticated commonsense reasoning, a capability naturally possessed by large-scale Vision-Language Models (VLMs). While VLMs show promise as zero-shot planners, their lack of grounded physical understanding…

Robotics · Computer Science 2026-03-18 Emily Yue-Ting Jia , Weiduo Yuan , Tianheng Shi , Vitor Guizilini , Jiageng Mao , Yue Wang

3D vision-language (VL) reasoning has gained significant attention due to its potential to bridge the 3D physical world with natural language descriptions. Existing approaches typically follow task-specific, highly specialized paradigms.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Hao Liu , Yanni Ma , Yan Liu , Haihong Xiao , Ying He

Vision-Language Models (VLM) can generate plausible high-level plans when prompted with a goal, the context, an image of the scene, and any planning constraints. However, there is no guarantee that the predicted actions are geometrically…

Robotics · Computer Science 2024-10-04 Zhutian Yang , Caelan Garrett , Dieter Fox , Tomás Lozano-Pérez , Leslie Pack Kaelbling

Implicit neural representations have demonstrated significant promise for 3D scene reconstruction. Recent works have extended their applications to autonomous implicit reconstruction through the Next Best View (NBV) based method. However,…

Robotics · Computer Science 2024-04-17 Jing Zeng , Yanxu Li , Jiahao Sun , Qi Ye , Yunlong Ran , Jiming Chen

Zero-shot Vision-and-Language Navigation (VLN) agents leveraging Large Language Models (LLMs) excel in generalization but suffer from insufficient spatial perception. Focusing on complex continuous environments, we categorize key perceptual…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Lu Yue , Yue Fan , Shiwei Lian , Yu Zhao , Jiaxin Yu , Liang Xie , Feitian Zhang