English
Related papers

Related papers: SpatialTrackerV2: 3D Point Tracking Made Easy

200 papers

Multi-Target Multi-Camera Tracking (MTMC) is an essential computer vision task for automating large-scale surveillance. With camera calibration and depth information, the targets in the scene can be projected into 3D space, offering…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Vu-Minh Le , Thao-Anh Tran , Duc Huy Do , Xuan Canh Do , Huong Ninh , Hai Tran

Direct methods have shown excellent performance in the applications of visual odometry and SLAM. In this work we propose to leverage their effectiveness for the task of 3D multi-object tracking. To this end, we propose DirectTracker, a…

Computer Vision and Pattern Recognition · Computer Science 2022-09-30 Mariia Gladkova , Nikita Korobov , Nikolaus Demmel , Aljoša Ošep , Laura Leal-Taixé , Daniel Cremers

Determining accurate bird's eye view (BEV) positions of objects and tracks in a scene is vital for various perception tasks including object interactions mapping, scenario extraction etc., however, the level of supervision required to…

Computer Vision and Pattern Recognition · Computer Science 2022-12-08 Paridhi Singh , Gaurav Singh , Arun Kumar

Reliable incremental estimation of camera poses and 3D reconstruction is key to enable various applications including robotics, interactive visualization, and augmented reality. However, this task is particularly challenging in dynamic…

Robotics · Computer Science 2025-12-09 Xingguang Zhong , Liren Jin , Marija Popović , Jens Behley , Cyrill Stachniss

The challenge of dynamic view synthesis from dynamic monocular videos, i.e., synthesizing novel views for free viewpoints given a monocular video of a dynamic scene captured by a moving camera, mainly lies in accurately modeling the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Meng You , Junhui Hou

Accurate and consistent 3D tracking from multiple cameras is a key component in a vision-based autonomous driving system. It involves modeling 3D dynamic objects in complex scenes across multiple cameras. This problem is inherently…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Tianyuan Zhang , Xuanyao Chen , Yue Wang , Yilun Wang , Hang Zhao

Point tracking aims to localize corresponding points across video frames, serving as a fundamental task for 4D reconstruction, robotics, and video editing. Existing methods commonly rely on shallow convolutional backbones such as ResNet…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Soowon Son , Honggyu An , Chaehyun Kim , Hyunah Ko , Jisu Nam , Dahyun Chung , Siyoon Jin , Jung Yi , Jaewon Min , Junhwa Hur , Seungryong Kim

We present the first approach to volumetric performance capture and novel-view rendering at real-time speed from monocular video, eliminating the need for expensive multi-view systems or cumbersome pre-acquisition of a personalized template…

Computer Vision and Pattern Recognition · Computer Science 2020-07-29 Ruilong Li , Yuliang Xiu , Shunsuke Saito , Zeng Huang , Kyle Olszewski , Hao Li

Template-based discriminative trackers are currently the dominant tracking paradigm due to their robustness, but are restricted to bounding box tracking and a limited range of transformation models, which reduces their localization…

Computer Vision and Pattern Recognition · Computer Science 2020-04-15 Alan Lukežič , Jiří Matas , Matej Kristan

Effective video tokenization is critical for scaling transformer models for long videos. Current approaches tokenize videos using space-time patches, leading to excessive tokens and computational inefficiencies. The best token reduction…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Chenhao Zheng , Jieyu Zhang , Mohammadreza Salehi , Ziqi Gao , Vishnu Iyengar , Norimasa Kobori , Quan Kong , Ranjay Krishna

Today's state-of-the-art methods for 3D object detection are based on lidar, stereo, or monocular cameras. Lidar-based methods achieve the best accuracy, but have a large footprint, high cost, and mechanically-limited angular sampling…

Computer Vision and Pattern Recognition · Computer Science 2021-02-09 Frank Julca-Aguilar , Jason Taylor , Mario Bijelic , Fahim Mannan , Ethan Tseng , Felix Heide

The field of indoor monocular 3D object detection is gaining significant attention, fueled by the increasing demand in VR/AR and robotic applications. However, its advancement is impeded by the limited availability and diversity of 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Jin-Cheng Jhang , Tao Tu , Fu-En Wang , Ke Zhang , Min Sun , Cheng-Hao Kuo

Reconstructing a dynamic target moving over a large area is challenging. Standard approaches for dynamic object reconstruction require dense coverage in both the viewing space and the temporal dimension, typically relying on multi-view…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Jun-Jee Chao , Volkan Isler

Traditional SLAM systems, which rely on bundle adjustment, struggle with highly dynamic scenes commonly found in casual videos. Such videos entangle the motion of dynamic elements, undermining the assumption of static environments required…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Weirong Chen , Ganlin Zhang , Felix Wimbauer , Rui Wang , Nikita Araslanov , Andrea Vedaldi , Daniel Cremers

To address the challenge of short-term object pose tracking in dynamic environments with monocular RGB input, we introduce a large-scale synthetic dataset OmniPose6D, crafted to mirror the diversity of real-world conditions. We additionally…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Yunzhi Lin , Yipu Zhao , Fu-Jen Chu , Xingyu Chen , Weiyao Wang , Hao Tang , Patricio A. Vela , Matt Feiszli , Kevin Liang

We present a near real-time method for 6-DoF tracking of an unknown object from a monocular RGBD video sequence, while simultaneously performing neural 3D reconstruction of the object. Our method works for arbitrary rigid objects, even when…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Bowen Wen , Jonathan Tremblay , Valts Blukis , Stephen Tyree , Thomas Muller , Alex Evans , Dieter Fox , Jan Kautz , Stan Birchfield

We propose an approach for reconstructing free-moving object from a monocular RGB video. Most existing methods either assume scene prior, hand pose prior, object category pose prior, or rely on local optimization with multiple sequence…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Haixin Shi , Yinlin Hu , Daniel Koguciuk , Juan-Ting Lin , Mathieu Salzmann , David Ferstl

We introduce a novel task of reconstructing a time series of second-person 3D human body meshes from monocular egocentric videos. The unique viewpoint and rapid embodied camera motion of egocentric videos raise additional technical barriers…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Miao Liu , Dexin Yang , Yan Zhang , Zhaopeng Cui , James M. Rehg , Siyu Tang

We present ARTrackV2, which integrates two pivotal aspects of tracking: determining where to look (localization) and how to describe (appearance analysis) the target object across video frames. Building on the foundation of its predecessor,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-14 Yifan Bai , Zeyang Zhao , Yihong Gong , Xing Wei

Recent advancements in zero-shot video diffusion models have shown promise for text-driven video editing, but challenges remain in achieving high temporal consistency. To address this, we introduce Video-3DGS, a 3D Gaussian Splatting…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Inkyu Shin , Qihang Yu , Xiaohui Shen , In So Kweon , Kuk-Jin Yoon , Liang-Chieh Chen
‹ Prev 1 8 9 10 Next ›