English
Related papers

Related papers: Grounding Image Matching in 3D with MASt3R

200 papers

We present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT)…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Yuanxun Lu , Jingyang Zhang , Tian Fang , Jean-Daniel Nahmias , Yanghai Tsin , Long Quan , Xun Cao , Yao Yao , Shiwei Li

Compared with the visual grounding on 2D images, the natural-language-guided 3D object localization on point clouds is more challenging. In this paper, we propose a new model, named InstanceRefer, to achieve a superior 3D visual grounding…

Computer Vision and Pattern Recognition · Computer Science 2021-07-30 Zhihao Yuan , Xu Yan , Yinghong Liao , Ruimao Zhang , Sheng Wang , Zhen Li , Shuguang Cui

This study addresses the challenge of performing visual localization in demanding conditions such as night-time scenarios, adverse weather, and seasonal changes. While many prior studies have focused on improving image-matching performance…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Khang Truong Giang , Soohwan Song , Sungho Jo

Recent learning methods for object pose estimation require resource-intensive training for each individual object instance or category, hampering their scalability in real applications when confronted with previously unseen objects. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Junwen Huang , Hao Yu , Kuan-Ting Yu , Nassir Navab , Slobodan Ilic , Benjamin Busam

Recent advances in 3D Gaussian Splatting (3DGS) present two main directions: feed-forward models offer fast inference in sparse-view settings, while per-scene optimization yields high-quality renderings but is computationally expensive. To…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Yueh-Cheng Liu , Jozef Hladký , Matthias Nießner , Angela Dai

Recovering dense 3D geometry from unposed images remains a foundational challenge in computer vision. Current state-of-the-art models are predominantly trained on perspective datasets, which implicitly constrains them to a standard pinhole…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Namitha Guruprasad , Abhay Yadav , Cheng Peng , Rama Chellappa

While recent feed-forward 3D reconstruction models accelerate 3D reconstruction by jointly inferring dense geometry and camera poses in a single pass, their reliance on dense attention imposes a quadratic complexity, creating a prohibitive…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Weining Ren , Xiao Tan , Kai Han

Recent advances in mapping techniques have enabled the creation of highly accurate dense 3D maps during robotic missions, such as point clouds, meshes, or NeRF-based representations. These developments present new opportunities for reusing…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Lintong Zhang , Yifu Tao , Jiarong Lin , Fu Zhang , Maurice Fallon

Robust 3D geometry estimation from videos is critical for applications such as autonomous navigation, SLAM, and 3D scene reconstruction. Recent methods like DUSt3R demonstrate that regressing dense pointmaps from image pairs enables…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Xiaoshan Wu , Yifei Yu , Xiaoyang Lyu , Yihua Huang , Bo Wang , Baoheng Zhang , Zhongrui Wang , Xiaojuan Qi

Rendering realistic images from 3D reconstruction is an essential task of many Computer Vision and Robotics pipelines, notably for mixed-reality applications as well as training autonomous agents in simulated environments. However, the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Lukas Bösiger , Mihai Dusmanu , Marc Pollefeys , Zuria Bauer

3D reconstruction in dynamic scenes primarily relies on the combination of geometry estimation and matching modules where the latter task is pivotal for distinguishing dynamic regions which can help to mitigate the interference introduced…

Image and Video Processing · Electrical Eng. & Systems 2025-08-11 Songyan Zhang , Yongtao Ge , Jinyuan Tian , Guangkai Xu , Hao Chen , Chen Lv , Chunhua Shen

The integration of aerial and ground images has been a promising solution in 3D modeling of complex scenes, which is seriously restricted by finding reliable correspondences. The primary contribution of this study is a feature matching…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Jiangxue Yu , Hui Wang , San Jiang , Xing Zhang , Dejin Zhang , Qingquan Li

We present a real-time monocular dense SLAM system designed bottom-up from MASt3R, a two-view 3D reconstruction and matching prior. Equipped with this strong prior, our system is robust on in-the-wild video sequences despite making no…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Riku Murai , Eric Dexheimer , Andrew J. Davison

Integration of aerial and ground images has been proved as an efficient approach to enhance the surface reconstruction in urban environments. However, as the first step, the feature point matching between aerial and ground images is…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Qing Zhu , Zhendong Wang , Han Hu , Linfu Xie , Xuming Ge , Yeting Zhang

Camera localization in 3D LiDAR maps has gained increasing attention due to its promising ability to handle complex scenarios, surpassing the limitations of visual-only localization methods. However, existing methods mostly focus on…

Robotics · Computer Science 2024-10-28 Huai Yu , Kuangyi Chen , Wen Yang , Sebastian Scherer , Gui-Song Xia

Room layout estimation from multiple-perspective images is poorly investigated due to the complexities that emerge from multi-view geometry, which requires muti-step solutions such as camera intrinsic and extrinsic estimation, image…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Yaxuan Huang , Xili Dai , Jianan Wang , Xianbiao Qi , Yixing Yuan , Xiangyu Yue

Erroneous feature matches have severe impact on subsequent camera pose estimation and often require additional, time-costly measures, like RANSAC, for outlier rejection. Our method tackles this challenge by addressing feature matching and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Barbara Roessle , Matthias Nießner

We address the problem of ground-to-satellite image geo-localization, that is, estimating the camera latitude, longitude and orientation (azimuth angle) by matching a query image captured at the ground level against a large-scale database…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Yujiao Shi , Xin Yu , Liu Liu , Dylan Campbell , Piotr Koniusz , Hongdong Li

Accurate and consistent 3D tracking from multiple cameras is a key component in a vision-based autonomous driving system. It involves modeling 3D dynamic objects in complex scenes across multiple cameras. This problem is inherently…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Tianyuan Zhang , Xuanyao Chen , Yue Wang , Yilun Wang , Hang Zhao

Object-centric scene understanding is a fundamental challenge in computer vision. Existing approaches often rely on multi-stage pipelines that first apply pre-trained segmentors to extract individual objects, followed by per-object 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Yi Du , Yang You , Xiang Wan , Leonidas Guibas