English
Related papers

Related papers: Bird's-Eye-View Scene Graph for Vision-Language Na…

200 papers

Aerial outdoor semantic navigation requires robots to explore large, unstructured environments to locate target objects. Recent advances in semantic navigation have demonstrated open-set object-goal navigation in indoor settings, but these…

World models have attracted increasing attention in autonomous driving for their ability to forecast potential future scenarios. In this paper, we propose BEVWorld, a novel framework that transforms multimodal sensor inputs into a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Yumeng Zhang , Shi Gong , Kaixin Xiong , Xiaoqing Ye , Xiaofan Li , Xiao Tan , Fan Wang , Jizhou Huang , Hua Wu , Haifeng Wang

Vision-and-Language Navigation (VLN) is a multi-modal, cooperative task requiring agents to interpret human instructions, navigate 3D environments, and communicate effectively under ambiguity. This paper presents a comprehensive review of…

Robotics · Computer Science 2025-12-02 Nivedan Yakolli , Avinash Gautam , Abhijit Das , Yuankai Qi , Virendra Singh Shekhawat

In recent years, 3D scene graphs have emerged as a powerful world representation, offering both geometric accuracy and semantic richness. Combining 3D scene graphs with large language models enables robots to reason, plan, and navigate in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Abdelrhman Werby , Dennis Rotondi , Fabio Scaparro , Kai O. Arras

We explore Bird's-Eye View (BEV) generation, converting a BEV map into its corresponding multi-view street images. Valued for its unified spatial representation aiding multi-sensor fusion, BEV is pivotal for various autonomous driving…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Xiaojie Xu , Tianshuo Xu , Fulong Ma , Yingcong Chen

Visual navigation in unknown environments based solely on natural language descriptions is a key capability for intelligent robots. In this work, we propose a navigation framework built upon off-the-shelf Visual Language Models (VLMs),…

Robotics · Computer Science 2025-08-08 Weifan Zhang , Tingguang Li , Yuzhen Liu

Current Visual Simultaneous Localization and Mapping (VSLAM) systems often struggle to create maps that are both semantically rich and easily interpretable. While incorporating semantic scene knowledge aids in building richer maps with…

LiDAR is crucial for robust 3D scene perception in autonomous driving. LiDAR perception has the largest body of literature after camera perception. However, multi-task learning across tasks like detection, segmentation, and motion…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Sambit Mohapatra , Senthil Yogamani , Varun Ravi Kumar , Stefan Milz , Heinrich Gotzig , Patrick Mäder

3D object detection in Bird's-Eye-View (BEV) space has recently emerged as a prevalent approach in the field of autonomous driving. Despite the demonstrated improvements in accuracy and velocity estimation compared to perspective view…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Yuxin Li , Qiang Han , Mengying Yu , Yuxin Jiang , Chaikiat Yeo , Yiheng Li , Zihang Huang , Nini Liu , Hsuanhan Chen , Xiaojun Wu

A 3D scene graph represents a compact scene model by capturing both the objects present and the semantic relationships between them, making it a promising structure for robotic applications. To effectively interact with users, an embodied…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Tatiana Zemskova , Dmitry Yudin

Vision-and-Language Navigation (VLN) tasks mainly evaluate agents based on one-time execution of individual instructions across multiple environments, aiming to develop agents capable of functioning in any environment in a zero-shot manner.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-30 Haodong Hong , Yanyuan Qiao , Sen Wang , Jiajun Liu , Qi Wu

Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through photorealistic environments following natural-language instructions. One main obstacle existing in VLN is data scarcity, leading to poor…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Yu Zhong , Rui Zhang , Zihao Zhang , Shuo Wang , Chuan Fang , Xishan Zhang , Jiaming Guo , Shaohui Peng , Di Huang , Yanyang Yan , Xing Hu , Qi Guo

Representing and understanding 3D environments in a structured manner is crucial for autonomous agents to navigate and reason about their surroundings. While traditional Simultaneous Localization and Mapping (SLAM) methods generate metric…

Robotics · Computer Science 2026-02-03 Albert Gassol Puigjaner , Angelos Zacharia , Kostas Alexis

Visual spatial description (VSD) aims to generate texts that describe the spatial relations of the given objects within images. Existing VSD work merely models the 2D geometrical vision features, thus inevitably falling prey to the problem…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Yu Zhao , Hao Fei , Wei Ji , Jianguo Wei , Meishan Zhang , Min Zhang , Tat-Seng Chua

Inspired by research in psychology, we introduce a behavioral approach for visual navigation using topological maps. Our goal is to enable a robot to navigate from one location to another, relying only on its visual input and the…

Computer Vision and Pattern Recognition · Computer Science 2019-03-04 Kevin Chen , Juan Pablo de Vicente , Gabriel Sepulveda , Fei Xia , Alvaro Soto , Marynel Vazquez , Silvio Savarese

Large Vision-Language Models (LVLMs) have achieved significant progress in tasks like visual question answering and document understanding. However, their potential to comprehend embodied environments and navigate within them remains…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Zhaowei Wang , Hongming Zhang , Tianqing Fang , Ye Tian , Yue Yang , Kaixin Ma , Xiaoman Pan , Yangqiu Song , Dong Yu

Autonomous navigation guided by natural language instructions in embodied environments remains a challenge for vision-language navigation (VLN) agents. Although recent advancements in learning diverse and fine-grained visual environmental…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Xuesong Zhang , Jia Li , Yunbo Xu , Zhenzhen Hu , Richang Hong

Using synthesized images to boost the performance of perception models is a long-standing research challenge in computer vision. It becomes more eminent in visual-centric autonomous driving systems with multi-view cameras as some long-tail…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Kairui Yang , Enhui Ma , Jibin Peng , Qing Guo , Di Lin , Kaicheng Yu

Visual Semantic Navigation (VSN) is a fundamental problem in robotics, where an agent must navigate toward a target object in an unknown environment, mainly using visual information. Most state-of-the-art VSN models are trained in…

While Open Set Semantic Mapping and 3D Semantic Scene Graphs (3DSSGs) are established paradigms in robotic perception, deploying them effectively to support high-level reasoning in large-scale, real-world environments remains a significant…

Robotics · Computer Science 2026-02-04 Martin Günther , Felix Igelbrink , Oscar Lima , Lennart Niecksch , Marian Renz , Martin Atzmueller