English
Related papers

Related papers: LoopNav: Benchmarking Spatial Consistency in World…

200 papers

Video understanding requires models to continuously track and update world state during playback. While existing benchmarks have advanced video understanding evaluation across multiple dimensions, the observation of how models maintain…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Pengyiang Liu , Zhongyue Shi , Hongye Hao , Qi Fu , Xueting Bi , Siwei Zhang , Xiaoyang Hu , Zitian Wang , Linjiang Huang , Si Liu

Visual robot navigation within large-scale, semi-structured environments deals with various challenges such as computation intensive path planning algorithms or insufficient knowledge about traversable spaces. Moreover, many…

Robotics · Computer Science 2018-03-12 Fabian Blöchliger , Marius Fehr , Marcin Dymczyk , Thomas Schneider , Roland Siegwart

Recent advances in deep monocular visual Simultaneous Localization and Mapping (SLAM) have achieved impressive accuracy and dense reconstruction capabilities, yet their robustness to scale inconsistency in large-scale indoor environments…

Robotics · Computer Science 2026-02-23 Hyoseok Ju , Bokeon Suh , Giseop Kim

Mapping is one of the crucial tasks enabling autonomous navigation of a mobile robot. Conventional mapping methods output a dense geometric map representation, e.g. an occupancy grid, which is not trivial to keep consistent for prolonged…

Robotics · Computer Science 2025-02-10 Kirill Muravyev , Alexander Melekhin , Dmitry Yudin , Konstantin Yakovlev

Semantic occupancy perception is essential for autonomous driving, as automated vehicles require a fine-grained perception of the 3D urban structures. However, existing relevant benchmarks lack diversity in urban scenes, and they only…

Computer Vision and Pattern Recognition · Computer Science 2023-03-08 Xiaofeng Wang , Zheng Zhu , Wenbo Xu , Yunpeng Zhang , Yi Wei , Xu Chi , Yun Ye , Dalong Du , Jiwen Lu , Xingang Wang

This paper proposes the SPARK dataset as a new unique space object multi-modal image dataset. Image-based object recognition is an important component of Space Situational Awareness, especially for applications such as on-orbit servicing,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-15 Mohamed Adel Musallam , Kassem Al Ismaeil , Oyebade Oyedotun , Marcos Damian Perez , Michel Poucet , Djamila Aouada

Visual navigation typically assumes the existence of at least one obstacle-free path between start and goal, which must be discovered/planned by the robot. However, in real-world scenarios, such as home environments and warehouses, clutter…

We study a challenging problem of unsupervised discovery of object landmarks. Many recent methods rely on bottlenecks to generate 2D Gaussian heatmaps however, these are limited in generating informed heatmaps while training, presumably due…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Mamona Awan , Muhammad Haris Khan , Sanoojan Baliah , Muhammad Ahmad Waseem , Salman Khan , Fahad Shahbaz Khan , Arif Mahmood

Generalization of imitation-learned navigation policies to environments unseen in training remains a major challenge. We address this by conducting the first large-scale study of how data quantity and data diversity affect real-world…

Recent advances in time-series forecasting increasingly rely on pre-trained foundation-style models. While these models often claim broad generalization, existing evaluation protocols provide limited evidence. Indeed, most current…

Where am I? This is one of the most critical questions that any intelligent system should answer to decide whether it navigates to a previously visited area. This problem has long been acknowledged for its challenging nature in simultaneous…

Robotics · Computer Science 2022-11-10 Konstantinos A. Tsintotas , Loukas Bampis , Antonios Gasteratos

Globally consistent dense maps are a key requirement for long-term robot navigation in complex environments. While previous works have addressed the challenges of dense mapping and global consistency, most require more computational…

In this paper, we compare different map management techniques for long-term visual navigation in changing environments. In this scenario, the navigation system needs to continuously update and refine its feature map in order to adapt to the…

Safe and explainable motion planning remains a central challenge in autonomous driving. While rule-based planners offer predictable and explainable behavior, they often fail to grasp the complexity and uncertainty of real-world traffic.…

Temporal consistency is critical in video prediction to ensure that outputs are coherent and free of artifacts. Traditional methods, such as temporal attention and 3D convolution, may struggle with significant object motion and may not…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Zihang Lai , Andrea Vedaldi

One of the main challenges in the Simultaneous Localization and Mapping (SLAM) loop closure problem is the recognition of previously visited places. In this work, we tackle the two main problems of real-time SLAM systems: 1) loop closure…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Mohammad-Maher Nakshbandi , Ziad Sharawy , Sorin Grigorescu

The creation of a metric-semantic map, which encodes human-prior knowledge, represents a high-level abstraction of environments. However, constructing such a map poses challenges related to the fusion of multi-modal sensor data, the…

Robotics · Computer Science 2024-12-03 Jianhao Jiao , Ruoyu Geng , Yuanhang Li , Ren Xin , Bowen Yang , Jin Wu , Lujia Wang , Ming Liu , Rui Fan , Dimitrios Kanoulas

Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies trained on expensive, embodiment-specific robot data. While…

Robotics · Computer Science 2026-03-17 Jie Chen , Yuxin Cai , Yizhuo Wang , Ruofei Bai , Yuhong Cao , Jun Li , Yau Wei Yun , Guillaume Sartoretti

Embodied navigation demands comprehensive scene understanding and precise spatial reasoning. While image-text models excel at interpreting pixel-level color and lighting cues, 3D-text models capture volumetric structure and spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Haihong Hao , Mingfei Han , Changlin Li , Zhihui Li , Xiaojun Chang

We introduce Matrix-Game, an interactive world foundation model for controllable game world generation. Matrix-Game is trained using a two-stage pipeline that first performs large-scale unlabeled pretraining for environment understanding,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Yifan Zhang , Chunli Peng , Boyang Wang , Puyi Wang , Qingcheng Zhu , Fei Kang , Biao Jiang , Zedong Gao , Eric Li , Yang Liu , Yahui Zhou