English
Related papers

Related papers: PointSt3R: Point Tracking through 3D Grounded Corr…

200 papers

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

We introduce CoTracker, a transformer-based model that tracks a large number of 2D points in long video sequences. Differently from most existing approaches that track points independently, CoTracker tracks them jointly, accounting for…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Nikita Karaev , Ignacio Rocco , Benjamin Graham , Natalia Neverova , Andrea Vedaldi , Christian Rupprecht

Large-scale vision foundation models have demonstrated remarkable success across various tasks, underscoring their robust generalization capabilities. While their proficiency in two-view correspondence has been explored, their effectiveness…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Görkay Aydemir , Weidi Xie , Fatma Güney

3D reconstruction, which aims to recover the dense three-dimensional structure of a scene, is a cornerstone technology for numerous applications, including augmented/virtual reality, autonomous driving, and robotics. While traditional…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Wei Zhang , Yihang Wu , Songhua Li , Wenjie Ma , Xin Ma , Qiang Li , Qi Wang

We introduce LocoTrack, a highly accurate and efficient model designed for the task of tracking any point (TAP) across video sequences. Previous approaches in this task often rely on local 2D correlation maps to establish correspondences…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Seokju Cho , Jiahui Huang , Jisu Nam , Honggyu An , Seungryong Kim , Joon-Young Lee

Extracting point correspondences from two or more views of a scene is a fundamental computer vision problem with particular importance for relative camera pose estimation and structure-from-motion. Existing local feature matching…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Dominik A. Kloepfer , João F. Henriques , Dylan Campbell

Narrated instructional videos often show and describe manipulations of similar objects, e.g., repairing a particular model of a car or laptop. In this work we aim to reconstruct such objects and to localize associated narrations in 3D.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-13 Dimitri Zhukov , Ignacio Rocco , Ivan Laptev , Josef Sivic , Johannes L. Schönberger , Bugra Tekin , Marc Pollefeys

We propose a novel online, point-based 3D reconstruction method from posed monocular RGB videos. Our model maintains a global point cloud representation of the scene, continuously updating the features and 3D locations of points as new…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Chen Ziwen , Zexiang Xu , Li Fuxin

State-of-the-art 3D computer vision algorithms continue to advance in handling sparse, unordered image sets. Recently developed foundational models for 3D reconstruction, such as Dense and Unconstrained Stereo 3D Reconstruction (DUSt3R),…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Xinyi Wu , Steven Landgraf , Markus Ulrich , Rongjun Qin

Recent approaches have successfully focused on the segmentation of static reconstructions, thereby equipping downstream applications with semantic 3D understanding. However, the world in which we live is dynamic, characterized by numerous…

Robotics · Computer Science 2025-03-12 Tjark Behrens , René Zurbrügg , Marc Pollefeys , Zuria Bauer , Hermann Blum

We introduce a novel, end-to-end learnable, differentiable non-rigid tracker that enables state-of-the-art non-rigid reconstruction by a learned robust optimization. Given two input RGB-D frames of a non-rigidly moving object, we employ a…

Computer Vision and Pattern Recognition · Computer Science 2021-01-13 Aljaž Božič , Pablo Palafox , Michael Zollhöfer , Angela Dai , Justus Thies , Matthias Nießner

This paper addresses the long-standing challenge of reconstructing 3D structures from videos with dynamic content. Current approaches to this problem were not designed to operate on casual videos recorded by standard cameras or require a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-28 Yoni Kasten , Wuyue Lu , Haggai Maron

3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models…

Robotics · Computer Science 2026-05-05 Sizhe Yang , Linning Xu , Hao Li , Juncheng Mu , Jia Zeng , Dahua Lin , Jiangmiao Pang

We present Point2Pose, a model-free method for causal 6D pose tracking of multiple rigid objects from monocular RGB-D video. Initialized only from sparse image points on the objects to be tracked, our approach tracks multiple unseen objects…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Tzu-Yuan Lin , Ho Jae Lee , Kevin Doherty , Yonghyeon Lee , Sangbae Kim

Despite recent success in incorporating learning into point cloud registration, many works focus on learning feature descriptors and continue to rely on nearest-neighbor feature matching and outlier filtering through RANSAC to obtain the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Zi Jian Yew , Gim Hee Lee

We introduce EgoPoints, a benchmark for point tracking in egocentric videos. We annotate 4.7K challenging tracks in egocentric sequences. Compared to the popular TAP-Vid-DAVIS evaluation benchmark, we include 9x more points that go…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Ahmad Darkhalil , Rhodri Guerrier , Adam W. Harley , Dima Damen

Modern Recurrent Neural Networks have become a competitive architecture for 3D reconstruction due to their linear-time complexity. However, their performance degrades significantly when applied beyond the training context length, revealing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Xingyu Chen , Yue Chen , Yuliang Xiu , Andreas Geiger , Anpei Chen

Faithfully reconstructing textured meshes is crucial for many applications. Compared to text or image modalities, leveraging 3D colored point clouds as input (colored-PC-to-mesh) offers inherent advantages in comprehensively and precisely…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Qiao Yu , Xianzhi Li , Yuan Tang , Xu Han , Jinfeng Xu , Long Hu , Min Chen

Accurate registration of 2D imagery with point clouds is a key technology for image-LiDAR point cloud fusion, camera to laser scanner calibration and camera localization. Despite continuous improvements, automatic registration of 2D and 3D…

Computer Vision and Pattern Recognition · Computer Science 2019-12-13 Huai Yu , Weikun Zhen , Wen Yang , Sebastian Scherer

Dense 3D reconstruction and ego-motion estimation are key challenges in autonomous driving and robotics. Compared to the complex, multi-modal systems deployed today, multi-camera systems provide a simpler, low-cost alternative. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Aron Schmied , Tobias Fischer , Martin Danelljan , Marc Pollefeys , Fisher Yu