中文
相关论文

相关论文: Cross-View Image Set Geo-Localization

200 篇论文

We introduce GSVisLoc, a visual localization method designed for 3D Gaussian Splatting (3DGS) scene representations. Given a 3DGS model of a scene and a query image, our goal is to estimate the camera's position and orientation. We…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Fadi Khatib , Dror Moran , Guy Trostianetsky , Yoni Kasten , Meirav Galun , Ronen Basri

Predicting the geographic location (geo-localization) from a single ground-level RGB image taken anywhere in the world is a very challenging problem. The challenges include huge diversity of images due to different environmental scenarios,…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Shraman Pramanick , Ewa M. Nowara , Joshua Gleason , Carlos D. Castillo , Rama Chellappa

The model-based gait recognition methods usually adopt the pedestrian walking postures to identify human beings. However, existing methods did not explicitly resolve the large intra-class variance of human pose due to camera views changing.…

计算机视觉与模式识别 · 计算机科学 2022-09-26 Honghu Pan , Yongyong Chen , Tingyang Xu , Yunqi He , Zhenyu He

3D Visual Grounding (3DVG) seeks to locate target objects in 3D scenes using natural language descriptions, enabling downstream applications such as augmented reality and robotics. Existing approaches typically rely on labeled 3D data and…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Rong Li , Shijie Li , Lingdong Kong , Xulei Yang , Junwei Liang

Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems to understand 3D environments and manipulate objects across…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Suchae Jeong , Jaehwi Song , Haeone Lee , Hanna Kim , Jian Kim , Dongjun Lee , Dong Kyu Shin , Changyeon Kim , Dongyoon Hahm , Woogyeol Jin , Juheon Choi , Kimin Lee

Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding performance by training large models with large-scale…

计算机视觉与模式识别 · 计算机科学 2025-09-25 Yangxiao Lu , Ruosen Li , Liqiang Jing , Jikai Wang , Xinya Du , Yunhui Guo , Nicholas Ruozzi , Yu Xiang

Video geolocalization is a crucial problem in current times. Given just a video, ascertaining where it was captured from can have a plethora of advantages. The problem of worldwide geolocalization has been tackled before, but only using the…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Parth Parag Kulkarni , Gaurav Kumar Nayak , Mubarak Shah

3D visual grounding (3DVG) involves localizing entities in a 3D scene referred to by natural language text. Such models are useful for embodied AI and scene retrieval applications, which involve searching for objects or patterns using…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Austin T. Wang , ZeMing Gong , Angel X. Chang

We propose an accurate and interpretable fine-grained cross-view localization method that estimates the 3 Degrees of Freedom (DoF) pose of a ground-level image by matching its local features with a reference aerial image. Unlike prior…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Zimin Xia , Chenghao Xu , Alexandre Alahi

We present a novel multi-altitude camera pose estimation system, addressing the challenges of robust and accurate localization across varied altitudes when only considering sparse image input. The system effectively handles diverse…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Yaxuan Li , Yewei Huang , Bijay Gaudel , Hamidreza Jafarnejadsani , Brendan Englot

Image geolocalization is the task of identifying the location depicted in a photo based only on its visual information. This task is inherently challenging since many photos have only few, possibly ambiguous cues to their geolocation.…

计算机视觉与模式识别 · 计算机科学 2018-08-08 Paul Hongsuck Seo , Tobias Weyand , Jack Sim , Bohyung Han

Federated learning (FL) has emerged as a promising paradigm for privacy-preserving multi-camera video understanding. However, applying FL to cross-view scenarios faces three major challenges: (i) heterogeneous viewpoints and backgrounds…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Shenghan Zhang , Run Ling , Ke Cao , Ao Ma , Zhanjie Zhang

The core problem of visual multi-robot simultaneous localization and mapping (MR-SLAM) is how to efficiently and accurately perform multi-robot global localization (MR-GL). The difficulties are two-fold. The first is the difficulty of…

机器人学 · 计算机科学 2021-02-25 Xiyue Guo , Junjie Hu , Junfeng Chen , Fuqin Deng , Tin Lun Lam

Recently, large vision-language models (LVLMs) unleash powerful analysis capabilities for low Earth orbit (LEO) satellite Earth observation images in the data center. However, fast satellite motion, brief satellite-ground station (GS)…

网络与互联网体系结构 · 计算机科学 2025-07-09 Yuxin Zhang , Jiahao Yang , Zhe Chen , Wenjun Zhu , Jin Zhao , Yue Gao

Multimodal intelligence development recently show strong progress in visual understanding and high level reasoning. Though, most reasoning system still reply on textual information as the main medium for inference. This limit their…

机器学习 · 计算机科学 2026-01-01 Soham Pahari , M. Srinivas

Determining the exact latitude and longitude that a photo was taken is a useful and widely applicable task, yet it remains exceptionally difficult despite the accelerated progress of other computer vision tasks. Most previous approaches…

计算机视觉与模式识别 · 计算机科学 2023-03-09 Brandon Clark , Alec Kerrigan , Parth Parag Kulkarni , Vicente Vivanco Cepeda , Mubarak Shah

Remote Sensing Visual Grounding (RSVG) aims to localize target objects in large-scale aerial imagery based on natural language descriptions. Owing to the vast spatial scale and high semantic ambiguity of remote sensing scenes, these…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Shiqi Huang , Shuting He , Bihan Wen

We propose a vision-based method that localizes a ground vehicle using publicly available satellite imagery as the only prior knowledge of the environment. Our approach takes as input a sequence of ground-level images acquired by the…

机器人学 · 计算机科学 2022-03-08 Dong-Ki Kim , Matthew R. Walter

The standard approach for visual place recognition is to use global image descriptors to retrieve the most similar database images for a given query image. The results can then be further improved with re-ranking methods that re-order the…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Gustav Hanning , Gabrielle Flood , Viktor Larsson

Vision-and-Language Navigation (VLN), where an agent follows instructions to reach a target destination, has recently seen significant advancements. In contrast to navigation in discrete environments with predefined trajectories, VLN in…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Guangzhao Dai , Jian Zhao , Yuantao Chen , Yusen Qin , Hao Zhao , Guosen Xie , Yazhou Yao , Xiangbo Shu , Xuelong Li
‹ 上一页 1 8 9 10 下一页 ›