中文
相关论文

相关论文: GMMLoc: Structure Consistent Visual Localization w…

200 篇论文

Worldwide geolocalization aims to locate the precise location at the coordinate level of photos taken anywhere on the Earth. It is very challenging due to 1) the difficulty of capturing subtle location-aware visual semantics, and 2) the…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Pengyue Jia , Yiding Liu , Xiaopeng Li , Yuhao Wang , Yantong Du , Xiao Han , Xuetao Wei , Shuaiqiang Wang , Dawei Yin , Xiangyu Zhao

Camera relocalization, a cornerstone capability of modern computer vision, accurately determines a camera's position and orientation (6-DoF) from images and is essential for applications in augmented reality (AR), mixed reality (MR),…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Zhendong Xiao , Wu Wei , Shujie Ji , Shan Yang , Changhao Chen

Cross-View Geo-Localization (CVGL) in remote sensing aims to locate a drone-view query by matching it to geo-tagged satellite images. Although supervised methods have achieved strong results on closeset benchmarks, they often fail to…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Jun Lu , Zehao Sang , Haoqi Wei , Xiangyun Liu , Kun Zhu , Haitao Guo , Zhihui Gong , Lei Ding

Visual localization plays an important role in the applications of Augmented Reality (AR), which enable AR devices to obtain their 6-DoF pose in the pre-build map in order to render virtual content in real scenes. However, most existing…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Hongjia Zhai , Xiyu Zhang , Boming Zhao , Hai Li , Yijia He , Zhaopeng Cui , Hujun Bao , Guofeng Zhang

Tracking the 6DoF pose of unknown objects in monocular RGB video sequences is crucial for robotic manipulation. However, existing approaches typically rely on accurate depth information, which is non-trivial to obtain in real-world…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Zhiyuan Chen , Fan Lu , Guo Yu , Bin Li , Sanqing Qu , Yuan Huang , Changhong Fu , Guang Chen

Since a building's floorplans are easily accessible, consistent over time, and inherently robust to changes in visual appearance, self-localization within the floorplan has attracted researchers' interest. However, since floorplans are…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Bolei Chen , Jiaxu Kang , Haonan Yang , Ping Zhong , Jianxin Wang

3D Gaussian Splatting (3DGS) has shown promising results for 3D scene modeling using mixtures of Gaussians, yet its existing simultaneous localization and mapping (SLAM) variants typically rely on direct, deterministic pose optimization…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Yuhan Zhu , Yanyu Zhang , Jie Xu , Wei Ren

Contextual optimization enhances decision quality by leveraging side information to improve predictions of uncertain parameters. However, existing approaches face significant challenges when dealing with multimodal or mixtures of…

最优化与控制 · 数学 2025-09-19 YoungChul Yoon , Grani A. Hanasusanto , Yijie Wang

This paper presents a novel camera relocalization method, STDLoc, which leverages Feature Gaussian as scene representation. STDLoc is a full relocalization pipeline that can achieve accurate relocalization without relying on any pose prior.…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Zhiwei Huang , Hailin Yu , Yichun Shentu , Jin Yuan , Guofeng Zhang

For reliable operation on urban roads, navigation using the Global Navigation Satellite System (GNSS) requires both accurately estimating the positioning detail from GNSS pseudorange measurements and determining when the estimated position…

机器人学 · 计算机科学 2021-10-26 Shubh Gupta , Grace X. Gao

Fast appearance variations and the distractions of similar objects are two of the most challenging problems in visual object tracking. Unlike many existing trackers that focus on modeling only the target, in this work, we consider the…

计算机视觉与模式识别 · 计算机科学 2020-08-28 Bi Li , Chengquan Zhang , Zhibin Hong , Xu Tang , Jingtuo Liu , Junyu Han , Errui Ding , Wenyu Liu

The visual simultaneous localization and mapping(vSLAM) is widely used in GPS-denied and open field environments for ground and surface robots. However, due to the frequent perception failures derived from lacking visual texture or the…

机器人学 · 计算机科学 2023-05-23 Zhihao Wang , Haoyao Chen , Shiwu Zhang , Yunjiang Lou

Visual localization tackles the challenge of estimating the camera pose from images by using correspondence analysis between query images and a map. This task is computation and data intensive which poses challenges on thorough evaluation…

Size, weight, and power constrained platforms impose constraints on computational resources that introduce unique challenges in implementing localization algorithms. We present a framework to perform fast localization on such platforms…

机器人学 · 计算机科学 2018-04-02 Aditya Dhawale , Kumar Shaurya Shankar , Nathan Michael

The precise prediction of human mobility has produced significant socioeconomic impacts, such as location recommendations and evacuation suggestions. However, existing methods suffer from limited generalization capability: unimodal…

人工智能 · 计算机科学 2025-12-30 Junshu Dai , Yu Wang , Tongya Zheng , Wei Ji , Qinghong Guo , Ji Cao , Jie Song , Canghong Jin , Mingli Song

This paper presents a hybrid real-time camera pose estimation framework with a novel partitioning scheme and introduces motion averaging to monocular Simultaneous Localization and Mapping (SLAM) systems. Breaking through the limitations of…

计算机视觉与模式识别 · 计算机科学 2020-11-04 Xinyi Li , Haibin Ling

At modern construction sites, utilizing GNSS (Global Navigation Satellite System) to measure the real-time location and orientation (i.e. pose) of construction machines and navigate them is very common. However, GNSS is not always…

机器人学 · 计算机科学 2021-01-19 Runqiu Bao , Ren Komatsu , Renato Miyagusuku , Masaki Chino , Atsushi Yamashita , Hajime Asama

Cross-view geo-localization aims to estimate the location of a query ground image by matching it to a reference geo-tagged aerial images database. As an extremely challenging task, its difficulties root in the drastic view changes and…

计算机视觉与模式识别 · 计算机科学 2023-06-19 Xiaohan Zhang , Xingyu Li , Waqas Sultani , Yi Zhou , Safwan Wshah

Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tasks like visual question answering (VQA), image captioning,…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Junxiao Xue , Quan Deng , Fei Yu , Yanhao Wang , Jun Wang , Yuehua Li

3D semantic occupancy prediction is a pivotal task in autonomous driving, providing a dense and fine-grained understanding of the surrounding environment, yet single-modality methods face trade-offs between camera semantics and LiDAR…

计算机视觉与模式识别 · 计算机科学 2026-02-02 A. Enes Doruk , Hasan F. Ates