中文
相关论文

相关论文: From 2D to 3D: AISG-SLA Visual Localization Challe…

200 篇论文

This report presents our team's 'PCIE_LAM' solution for the Ego4D Looking At Me Challenge at CVPR2024. The main goal of the challenge is to accurately determine if a person in the scene is looking at the camera wearer, based on a video…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Kanokphan Lertniphonphan , Jun Xie , Yaqing Meng , Shijing Wang , Feng Chen , Zhepeng Wang

The field of visual localization has been researched for several decades and has meanwhile found many practical applications. Despite the strong progress in this field, there are still challenging situations in which established methods…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Vincent Ress , Jonas Meyer , Wei Zhang , David Skuddis , Uwe Soergel , Norbert Haala

Vision-based localization for autonomous driving has been of great interest among researchers. When a pre-built 3D map is not available, the techniques of visual simultaneous localization and mapping (SLAM) are typically adopted. Due to…

机器人学 · 计算机科学 2024-04-16 Yanhao Zhang , Yujiao Shi , Shan Wang , Ankit Vora , Akhil Perincherry , Yongbo Chen , Hongdong Li

This paper presents a novel system designed for 3D mapping and visual relocalization using 3D Gaussian Splatting. Our proposed method uses LiDAR and camera data to create accurate and visually plausible representations of the environment.…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Peng Jiang , Gaurav Pandey , Srikanth Saripalli

Simultaneous localization and mapping (SLAM) has achieved impressive performance in static environments. However, SLAM in dynamic environments remains an open question. Many methods directly filter out dynamic objects, resulting in…

机器人学 · 计算机科学 2024-11-26 Haoang Li , Xiangqi Meng , Xingxing Zuo , Zhe Liu , Hesheng Wang , Daniel Cremers

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Meng Wang , Huilong Pi , Ruihui Li , Yunchuan Qin , Zhuo Tang , Kenli Li

The Imagenet Large Scale Visual Recognition Challenge (ILSVRC) is the one of the most important big data challenges to date. We participated in the object detection track of ILSVRC 2014 and received the fourth place among the 38 teams. We…

计算机视觉与模式识别 · 计算机科学 2014-10-07 Cewu Lu , Hao Chen , Qifeng Chen , Hei Law , Yao Xiao , Chi-Keung Tang

Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring cross-viewpoint understanding and spatial reasoning. We…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Dingming Li , Hongxing Li , Zixuan Wang , Yuchen Yan , Hang Zhang , Siqi Chen , Guiyang Hou , Shengpei Jiang , Wenqi Zhang , Yongliang Shen , Weiming Lu , Yueting Zhuang

Topological localization is a fundamental problem in mobile robotics, since robots must be able to determine their position in order to accomplish tasks. Visual localization and place recognition are challenging due to perceptual ambiguity,…

机器人学 · 计算机科学 2025-09-08 Emanuela Boros

In vision-and-language navigation (VLN), an embodied agent is required to navigate in realistic 3D environments following natural language instructions. One major bottleneck for existing VLN approaches is the lack of sufficient training…

计算机视觉与模式识别 · 计算机科学 2022-08-26 Shizhe Chen , Pierre-Louis Guhur , Makarand Tapaswi , Cordelia Schmid , Ivan Laptev

We present VLPG-Nav, a visual language navigation method for guiding robots to specified objects within household scenes. Unlike existing methods primarily focused on navigating the robot toward objects, our approach considers the…

The forthcoming Fifth Generation (5G) era raises the expectation for ubiquitous wireless connectivity to enhance human experiences in information and knowledge sharing as well as in entertainment and social interactions. The promising…

信息论 · 计算机科学 2017-09-07 Rong Zhang

Modern deep learning developments create new opportunities for 3D mapping technology, scene reconstruction pipelines, and virtual reality development. Despite advances in 3D deep learning technology, direct training of deep learning models…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Xueyang Kang

Perceiving humans in the context of Intelligent Transportation Systems (ITS) often relies on multiple cameras or expensive LiDAR sensors. In this work, we present a new cost-effective vision-based method that perceives humans' locations in…

计算机视觉与模式识别 · 计算机科学 2021-05-04 Lorenzo Bertoni , Sven Kreiss , Alexandre Alahi

In this paper, we propose a method for initial camera pose estimation from just a single image which is robust to viewing conditions and does not require a detailed model of the scene. This method meets the growing need of easy deployment…

计算机视觉与模式识别 · 计算机科学 2022-03-10 Matthieu Zins , Gilles Simon , Marie-Odile Berger

Many robotics applications require precise pose estimates despite operating in large and changing environments. This can be addressed by visual localization, using a pre-computed 3D model of the surroundings. The pose estimation then…

计算机视觉与模式识别 · 计算机科学 2018-09-20 Paul-Edouard Sarlin , Frédéric Debraine , Marcin Dymczyk , Roland Siegwart , Cesar Cadena

Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global audio and one-scale…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Shentong Mo , Haofan Wang

Visual relocalization has been a widely discussed problem in 3D vision: given a pre-constructed 3D visual map, the 6 DoF (Degrees-of-Freedom) pose of a query image is estimated. Relocalization in large-scale indoor environments enables…

计算机视觉与模式识别 · 计算机科学 2022-07-27 Jiahui Zhang , Shitao Tang , Kejie Qiu , Rui Huang , Chuan Fang , Le Cui , Zilong Dong , Siyu Zhu , Ping Tan

Place recognition is a core component of Simultaneous Localization and Mapping (SLAM) algorithms. Particularly in visual SLAM systems, previously-visited places are recognized by measuring the appearance similarity between images…

计算机视觉与模式识别 · 计算机科学 2020-07-28 Jiawei Mo , Junaed Sattar

Visual queries 3D localization (VQ3D) is a task in the Ego4D Episodic Memory Benchmark. Given an egocentric video, the goal is to answer queries of the form "Where did I last see object X?", where the query object X is specified as a static…

计算机视觉与模式识别 · 计算机科学 2022-11-21 Jinjie Mai , Chen Zhao , Abdullah Hamdi , Silvio Giancola , Bernard Ghanem