中文
相关论文

相关论文: $L^3$:Scene-agnostic Visual Localization in the Wi…

200 篇论文

Image geolocation is a critical task in various image-understanding applications. However, existing methods often fail when analyzing challenging, in-the-wild images. Inspired by the exceptional background knowledge of multimodal language…

计算机视觉与模式识别 · 计算机科学 2024-06-03 Zhiqiang Wang , Dejia Xu , Rana Muhammad Shahroz Khan , Yanbin Lin , Zhiwen Fan , Xingquan Zhu

Modern 3D object detection datasets are constrained by narrow class taxonomies and costly manual annotations, limiting their ability to scale to open-world settings. In contrast, 2D vision-language models trained on web-scale image-text…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Atharv Goel , Mehar Khurana

We are interested in automatic scene understanding from geometric cues. To this end, we aim to bring semantic segmentation in the loop of real-time reconstruction. Our semantic segmentation is built on a deep autoencoder stack trained…

计算机视觉与模式识别 · 计算机科学 2015-05-04 Ankur Handa , Viorica Patraucean , Vijay Badrinarayanan , Simon Stent , Roberto Cipolla

We propose Flash3D, a method for scene reconstruction and novel view synthesis from a single image which is both very generalisable and efficient. For generalisability, we start from a "foundation" model for monocular depth estimation and…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Stanislaw Szymanowicz , Eldar Insafutdinov , Chuanxia Zheng , Dylan Campbell , João F. Henriques , Christian Rupprecht , Andrea Vedaldi

Accurate and robust visual localization under a wide range of viewing condition variations including season and illumination changes, as well as weather and day-night variations, is the key component for many computer vision and robotics…

计算机视觉与模式识别 · 计算机科学 2019-05-20 Tianxin Shi , Shuhan Shen , Xiang Gao , Lingjie Zhu

Full autonomy for fixed-wing unmanned aerial vehicles (UAVs) requires the capability to autonomously detect potential landing sites in unknown and unstructured terrain, allowing for self-governed mission completion or handling of emergency…

机器人学 · 计算机科学 2018-02-27 Timo Hinzmann , Thomas Stastny , Cesar Cadena , Roland Siegwart , Igor Gilitschenski

Retrieval-based place recognition is an efficient and effective solution for re-localization within a pre-built map, or global data association for Simultaneous Localization and Mapping (SLAM). The accuracy of such an approach is heavily…

计算机视觉与模式识别 · 计算机科学 2022-09-27 Kavisha Vidanapathirana , Milad Ramezani , Peyman Moghadam , Sridha Sridharan , Clinton Fookes

Abstract--- Exploiting the spatial structure in scene images is a key research direction for scene recognition. Due to the large intra-class structural diversity, building and modeling flexible structural layout to adapt various image…

计算机视觉与模式识别 · 计算机科学 2020-06-24 Gongwei Chen , Xinhang Song , Haitao Zeng , Shuqiang Jiang

We present a novel 3D pose refinement approach based on differentiable rendering for objects of arbitrary categories in the wild. In contrast to previous methods, we make two main contributions: First, instead of comparing real-world images…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Alexander Grabner , Yaming Wang , Peizhao Zhang , Peihong Guo , Tong Xiao , Peter Vajda , Peter M. Roth , Vincent Lepetit

Synthesizing interactive 3D scenes from text is essential for gaming, virtual reality, and embodied AI. However, existing methods face several challenges. Learning-based approaches depend on small-scale indoor datasets, limiting the scene…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Lu Ling , Chen-Hsuan Lin , Tsung-Yi Lin , Yifan Ding , Yu Zeng , Yichen Sheng , Yunhao Ge , Ming-Yu Liu , Aniket Bera , Zhaoshuo Li

We present LT3SD, a novel latent diffusion model for large-scale 3D scene generation. Recent advances in diffusion models have shown impressive results in 3D object generation, but are limited in spatial extent and quality when extended to…

计算机视觉与模式识别 · 计算机科学 2025-05-02 Quan Meng , Lei Li , Matthias Nießner , Angela Dai

The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environments. Efforts to…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Duo Zheng , Shijia Huang , Liwei Wang

In crowded urban environments where traffic is dense, current technologies struggle to oversee tight navigation, but surface-level understanding allows autonomous vehicles to safely assess proximity to surrounding obstacles. 3D or 2D scene…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Akarshani Ramanayake , Nihal Kodikara

We present LaLaLoc to localise in environments without the need for prior visitation, and in a manner that is robust to large changes in scene appearance, such as a full rearrangement of furniture. Specifically, LaLaLoc performs…

计算机视觉与模式识别 · 计算机科学 2021-10-13 Henry Howard-Jenkins , Jose-Raul Ruiz-Sarmiento , Victor Adrian Prisacariu

Visual localization is of great importance in robotics and computer vision. Recently, scene coordinate regression based methods have shown good performance in visual localization in small static scenes. However, it still estimates camera…

计算机视觉与模式识别 · 计算机科学 2021-05-25 Zhaoyang Huang , Han Zhou , Yijin Li , Bangbang Yang , Yan Xu , Xiaowei Zhou , Hujun Bao , Guofeng Zhang , Hongsheng Li

Global localization is an important and widely studied problem for many robotic applications. Place recognition approaches can be exploited to solve this task, e.g., in the autonomous driving field. While most vision-based approaches match…

计算机视觉与模式识别 · 计算机科学 2020-03-11 Daniele Cattaneo , Matteo Vaghi , Simone Fontana , Augusto Luis Ballardini , Domenico Giorgio Sorrenti

High-definition maps (HD maps) are a key component of most modern self-driving systems due to their valuable semantic and geometric information. Unfortunately, building HD maps has proven hard to scale due to their cost as well as the…

机器人学 · 计算机科学 2021-01-19 Sergio Casas , Abbas Sadat , Raquel Urtasun

Visual localization in large-scale UAV scenarios is a critical capability for autonomous systems, yet it remains challenging due to geometric complexity and environmental variations. While 3D Gaussian Splatting (3DGS) has emerged as a…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Xiang Zhang , Tengfei Wang , Fang Xu , Xin Wang , Zongqian Zhan

Reconstructing 3D scenes from sparse, unposed images remains challenging under real-world conditions with varying illumination and transient occlusions. Existing methods rely on scene-specific optimization using appearance embeddings or…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Vinayak Gupta , Chih-Hao Lin , Shenlong Wang , Anand Bhattad , Jia-Bin Huang

We introduce a data-driven approach for interactively synthesizing in-the-wild images from semantic label maps. Our approach is dramatically different from recent work in this space, in that we make use of no learning. Instead, our approach…

计算机视觉与模式识别 · 计算机科学 2019-06-12 Aayush Bansal , Yaser Sheikh , Deva Ramanan