English
Related papers

Related papers: From 2D to 3D: AISG-SLA Visual Localization Challe…

200 papers

This report presents our team's 'PCIE_LAM' solution for the Ego4D Looking At Me Challenge at CVPR2024. The main goal of the challenge is to accurately determine if a person in the scene is looking at the camera wearer, based on a video…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Kanokphan Lertniphonphan , Jun Xie , Yaqing Meng , Shijing Wang , Feng Chen , Zhepeng Wang

The field of visual localization has been researched for several decades and has meanwhile found many practical applications. Despite the strong progress in this field, there are still challenging situations in which established methods…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Vincent Ress , Jonas Meyer , Wei Zhang , David Skuddis , Uwe Soergel , Norbert Haala

Vision-based localization for autonomous driving has been of great interest among researchers. When a pre-built 3D map is not available, the techniques of visual simultaneous localization and mapping (SLAM) are typically adopted. Due to…

Robotics · Computer Science 2024-04-16 Yanhao Zhang , Yujiao Shi , Shan Wang , Ankit Vora , Akhil Perincherry , Yongbo Chen , Hongdong Li

This paper presents a novel system designed for 3D mapping and visual relocalization using 3D Gaussian Splatting. Our proposed method uses LiDAR and camera data to create accurate and visually plausible representations of the environment.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Peng Jiang , Gaurav Pandey , Srikanth Saripalli

Simultaneous localization and mapping (SLAM) has achieved impressive performance in static environments. However, SLAM in dynamic environments remains an open question. Many methods directly filter out dynamic objects, resulting in…

Robotics · Computer Science 2024-11-26 Haoang Li , Xiangqi Meng , Xingxing Zuo , Zhe Liu , Hesheng Wang , Daniel Cremers

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Meng Wang , Huilong Pi , Ruihui Li , Yunchuan Qin , Zhuo Tang , Kenli Li

The Imagenet Large Scale Visual Recognition Challenge (ILSVRC) is the one of the most important big data challenges to date. We participated in the object detection track of ILSVRC 2014 and received the fourth place among the 38 teams. We…

Computer Vision and Pattern Recognition · Computer Science 2014-10-07 Cewu Lu , Hao Chen , Qifeng Chen , Hei Law , Yao Xiao , Chi-Keung Tang

Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring cross-viewpoint understanding and spatial reasoning. We…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Dingming Li , Hongxing Li , Zixuan Wang , Yuchen Yan , Hang Zhang , Siqi Chen , Guiyang Hou , Shengpei Jiang , Wenqi Zhang , Yongliang Shen , Weiming Lu , Yueting Zhuang

Topological localization is a fundamental problem in mobile robotics, since robots must be able to determine their position in order to accomplish tasks. Visual localization and place recognition are challenging due to perceptual ambiguity,…

Robotics · Computer Science 2025-09-08 Emanuela Boros

In vision-and-language navigation (VLN), an embodied agent is required to navigate in realistic 3D environments following natural language instructions. One major bottleneck for existing VLN approaches is the lack of sufficient training…

Computer Vision and Pattern Recognition · Computer Science 2022-08-26 Shizhe Chen , Pierre-Louis Guhur , Makarand Tapaswi , Cordelia Schmid , Ivan Laptev

We present VLPG-Nav, a visual language navigation method for guiding robots to specified objects within household scenes. Unlike existing methods primarily focused on navigating the robot toward objects, our approach considers the…

The forthcoming Fifth Generation (5G) era raises the expectation for ubiquitous wireless connectivity to enhance human experiences in information and knowledge sharing as well as in entertainment and social interactions. The promising…

Information Theory · Computer Science 2017-09-07 Rong Zhang

Modern deep learning developments create new opportunities for 3D mapping technology, scene reconstruction pipelines, and virtual reality development. Despite advances in 3D deep learning technology, direct training of deep learning models…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Xueyang Kang

Perceiving humans in the context of Intelligent Transportation Systems (ITS) often relies on multiple cameras or expensive LiDAR sensors. In this work, we present a new cost-effective vision-based method that perceives humans' locations in…

Computer Vision and Pattern Recognition · Computer Science 2021-05-04 Lorenzo Bertoni , Sven Kreiss , Alexandre Alahi

In this paper, we propose a method for initial camera pose estimation from just a single image which is robust to viewing conditions and does not require a detailed model of the scene. This method meets the growing need of easy deployment…

Computer Vision and Pattern Recognition · Computer Science 2022-03-10 Matthieu Zins , Gilles Simon , Marie-Odile Berger

Many robotics applications require precise pose estimates despite operating in large and changing environments. This can be addressed by visual localization, using a pre-computed 3D model of the surroundings. The pose estimation then…

Computer Vision and Pattern Recognition · Computer Science 2018-09-20 Paul-Edouard Sarlin , Frédéric Debraine , Marcin Dymczyk , Roland Siegwart , Cesar Cadena

Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global audio and one-scale…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Shentong Mo , Haofan Wang

Visual relocalization has been a widely discussed problem in 3D vision: given a pre-constructed 3D visual map, the 6 DoF (Degrees-of-Freedom) pose of a query image is estimated. Relocalization in large-scale indoor environments enables…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Jiahui Zhang , Shitao Tang , Kejie Qiu , Rui Huang , Chuan Fang , Le Cui , Zilong Dong , Siyu Zhu , Ping Tan

Place recognition is a core component of Simultaneous Localization and Mapping (SLAM) algorithms. Particularly in visual SLAM systems, previously-visited places are recognized by measuring the appearance similarity between images…

Computer Vision and Pattern Recognition · Computer Science 2020-07-28 Jiawei Mo , Junaed Sattar

Visual queries 3D localization (VQ3D) is a task in the Ego4D Episodic Memory Benchmark. Given an egocentric video, the goal is to answer queries of the form "Where did I last see object X?", where the query object X is specified as a static…

Computer Vision and Pattern Recognition · Computer Science 2022-11-21 Jinjie Mai , Chen Zhao , Abdullah Hamdi , Silvio Giancola , Bernard Ghanem