中文
相关论文

相关论文: PointSt3R: Point Tracking through 3D Grounded Corr…

200 篇论文

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

We introduce CoTracker, a transformer-based model that tracks a large number of 2D points in long video sequences. Differently from most existing approaches that track points independently, CoTracker tracks them jointly, accounting for…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Nikita Karaev , Ignacio Rocco , Benjamin Graham , Natalia Neverova , Andrea Vedaldi , Christian Rupprecht

Large-scale vision foundation models have demonstrated remarkable success across various tasks, underscoring their robust generalization capabilities. While their proficiency in two-view correspondence has been explored, their effectiveness…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Görkay Aydemir , Weidi Xie , Fatma Güney

3D reconstruction, which aims to recover the dense three-dimensional structure of a scene, is a cornerstone technology for numerous applications, including augmented/virtual reality, autonomous driving, and robotics. While traditional…

计算机视觉与模式识别 · 计算机科学 2025-07-14 Wei Zhang , Yihang Wu , Songhua Li , Wenjie Ma , Xin Ma , Qiang Li , Qi Wang

We introduce LocoTrack, a highly accurate and efficient model designed for the task of tracking any point (TAP) across video sequences. Previous approaches in this task often rely on local 2D correlation maps to establish correspondences…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Seokju Cho , Jiahui Huang , Jisu Nam , Honggyu An , Seungryong Kim , Joon-Young Lee

Extracting point correspondences from two or more views of a scene is a fundamental computer vision problem with particular importance for relative camera pose estimation and structure-from-motion. Existing local feature matching…

计算机视觉与模式识别 · 计算机科学 2024-01-22 Dominik A. Kloepfer , João F. Henriques , Dylan Campbell

Narrated instructional videos often show and describe manipulations of similar objects, e.g., repairing a particular model of a car or laptop. In this work we aim to reconstruct such objects and to localize associated narrations in 3D.…

计算机视觉与模式识别 · 计算机科学 2021-09-13 Dimitri Zhukov , Ignacio Rocco , Ivan Laptev , Josef Sivic , Johannes L. Schönberger , Bugra Tekin , Marc Pollefeys

We propose a novel online, point-based 3D reconstruction method from posed monocular RGB videos. Our model maintains a global point cloud representation of the scene, continuously updating the features and 3D locations of points as new…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Chen Ziwen , Zexiang Xu , Li Fuxin

State-of-the-art 3D computer vision algorithms continue to advance in handling sparse, unordered image sets. Recently developed foundational models for 3D reconstruction, such as Dense and Unconstrained Stereo 3D Reconstruction (DUSt3R),…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Xinyi Wu , Steven Landgraf , Markus Ulrich , Rongjun Qin

Recent approaches have successfully focused on the segmentation of static reconstructions, thereby equipping downstream applications with semantic 3D understanding. However, the world in which we live is dynamic, characterized by numerous…

机器人学 · 计算机科学 2025-03-12 Tjark Behrens , René Zurbrügg , Marc Pollefeys , Zuria Bauer , Hermann Blum

We introduce a novel, end-to-end learnable, differentiable non-rigid tracker that enables state-of-the-art non-rigid reconstruction by a learned robust optimization. Given two input RGB-D frames of a non-rigidly moving object, we employ a…

计算机视觉与模式识别 · 计算机科学 2021-01-13 Aljaž Božič , Pablo Palafox , Michael Zollhöfer , Angela Dai , Justus Thies , Matthias Nießner

This paper addresses the long-standing challenge of reconstructing 3D structures from videos with dynamic content. Current approaches to this problem were not designed to operate on casual videos recorded by standard cameras or require a…

计算机视觉与模式识别 · 计算机科学 2024-06-28 Yoni Kasten , Wuyue Lu , Haggai Maron

3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models…

机器人学 · 计算机科学 2026-05-05 Sizhe Yang , Linning Xu , Hao Li , Juncheng Mu , Jia Zeng , Dahua Lin , Jiangmiao Pang

We present Point2Pose, a model-free method for causal 6D pose tracking of multiple rigid objects from monocular RGB-D video. Initialized only from sparse image points on the objects to be tracked, our approach tracks multiple unseen objects…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Tzu-Yuan Lin , Ho Jae Lee , Kevin Doherty , Yonghyeon Lee , Sangbae Kim

Despite recent success in incorporating learning into point cloud registration, many works focus on learning feature descriptors and continue to rely on nearest-neighbor feature matching and outlier filtering through RANSAC to obtain the…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Zi Jian Yew , Gim Hee Lee

We introduce EgoPoints, a benchmark for point tracking in egocentric videos. We annotate 4.7K challenging tracks in egocentric sequences. Compared to the popular TAP-Vid-DAVIS evaluation benchmark, we include 9x more points that go…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Ahmad Darkhalil , Rhodri Guerrier , Adam W. Harley , Dima Damen

Modern Recurrent Neural Networks have become a competitive architecture for 3D reconstruction due to their linear-time complexity. However, their performance degrades significantly when applied beyond the training context length, revealing…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Xingyu Chen , Yue Chen , Yuliang Xiu , Andreas Geiger , Anpei Chen

Faithfully reconstructing textured meshes is crucial for many applications. Compared to text or image modalities, leveraging 3D colored point clouds as input (colored-PC-to-mesh) offers inherent advantages in comprehensively and precisely…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Qiao Yu , Xianzhi Li , Yuan Tang , Xu Han , Jinfeng Xu , Long Hu , Min Chen

Accurate registration of 2D imagery with point clouds is a key technology for image-LiDAR point cloud fusion, camera to laser scanner calibration and camera localization. Despite continuous improvements, automatic registration of 2D and 3D…

计算机视觉与模式识别 · 计算机科学 2019-12-13 Huai Yu , Weikun Zhen , Wen Yang , Sebastian Scherer

Dense 3D reconstruction and ego-motion estimation are key challenges in autonomous driving and robotics. Compared to the complex, multi-modal systems deployed today, multi-camera systems provide a simpler, low-cost alternative. However,…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Aron Schmied , Tobias Fischer , Martin Danelljan , Marc Pollefeys , Fisher Yu