English
Related papers

Related papers: Track4World: Feedforward World-centric Dense 3D Tr…

200 papers

We present a fully automatic approach to real-time 3D face reconstruction from monocular in-the-wild videos. With the use of a cascaded-regressor based face tracking and a 3D Morphable Face Model shape fitting, we obtain a semi-dense 3D…

Computer Vision and Pattern Recognition · Computer Science 2017-08-28 Patrik Huber , Philipp Kopp , Matthias Rätsch , William Christmas , Josef Kittler

Closed-loop simulation is essential for advancing end-to-end autonomous driving systems. Contemporary sensor simulation methods, such as NeRF and 3DGS, rely predominantly on conditions closely aligned with training data distributions, which…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Guosheng Zhao , Chaojun Ni , Xiaofeng Wang , Zheng Zhu , Xueyang Zhang , Yida Wang , Guan Huang , Xinze Chen , Boyuan Wang , Youyi Zhang , Wenjun Mei , Xingang Wang

We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction…

Creating a photorealistic scene and human reconstruction from a single monocular in-the-wild video figures prominently in the perception of a human-centric 3D world. Recent neural rendering advances have enabled holistic human-scene…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Zetong Zhang , Manuel Kaufmann , Lixin Xue , Jie Song , Martin R. Oswald

Multi-object tracking (MOT) in monocular videos is fundamentally challenged by occlusions and depth ambiguity, issues that conventional tracking-by-detection (TBD) methods struggle to resolve owing to a lack of geometric awareness. To…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xudong Han , Pengcheng Fang , Yueying Tian , Jianhui Yu , Xiaohao Cai , Daniel Roggen , Philip Birch

Recovering a dynamic 3D scene from a long monocular video is crucial for dense geometry, camera motion, and temporal correspondence to remain consistent in a shared coordinate system. Existing methods face two key challenges: (1)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Chenyi Xu , Yihao Wu , Liqi Yan , Chao Yang , Jianhui Zhang , Fangli Guan , Pan Li

A reliable and accurate 3D tracking framework is essential for predicting future locations of surrounding objects and planning the observer's actions in numerous applications such as autonomous driving. We propose a framework that can…

Computer Vision and Pattern Recognition · Computer Science 2021-03-15 Hou-Ning Hu , Yung-Hsu Yang , Tobias Fischer , Trevor Darrell , Fisher Yu , Min Sun

Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Melonie de Almeida , Daniela Ivanova , Tong Shi , John H. Williamson , Paul Henderson

Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Linyi Jin , Richard Tucker , Zhengqi Li , David Fouhey , Noah Snavely , Aleksander Holynski

We present a novel framework for dynamic radiance field prediction given monocular video streams. Unlike previous methods that primarily focus on predicting future frames, our method goes a step further by generating explicit 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Di Qi , Tong Yang , Beining Wang , Xiangyu Zhang , Wenqiang Zhang

We propose ProTracker, a novel framework for accurate and robust long-term dense tracking of arbitrary points in videos. Previous methods relying on global cost volumes effectively handle large occlusions and scene changes but lack…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Tingyang Zhang , Chen Wang , Zhiyang Dou , Qingzhe Gao , Jiahui Lei , Baoquan Chen , Lingjie Liu

Visual tracking has made significant improvements in the past few decades. Most existing state-of-the-art trackers 1) merely aim for performance in ideal conditions while overlooking the real-world conditions; 2) adopt the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Ziang Cao , Ziyuan Huang , Liang Pan , Shiwei Zhang , Ziwei Liu , Changhong Fu

While recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across frames, breaking visual coherence. We hypothesize that this…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Hyeonho Jeong , Chun-Hao Paul Huang , Jong Chul Ye , Niloy Mitra , Duygu Ceylan

Recent approaches to point tracking are able to recover the trajectory of any scene point through a large portion of a video despite the presence of occlusions. They are, however, too slow in practice to track every point observed in a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Guillaume Le Moing , Jean Ponce , Cordelia Schmid

While separately leveraging monocular 3D object detection and 2D multi-object tracking can be straightforwardly applied to sequence images in a frame-by-frame fashion, stand-alone tracker cuts off the transmission of the uncertainty from…

Computer Vision and Pattern Recognition · Computer Science 2022-05-31 Peixuan Li , Jieyu Jin

Predicting future motion is crucial in video understanding and controllable video generation. Dense point trajectories are a compact, expressive motion representation, but modeling their future evolution from observed video remains…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Zewei Zhang , Jia Jun Cheng Xian , Kaiwen Liu , Ming Liang , Hang Chu , Jun Chen , Renjie Liao

View-predictive generative models provide strong priors for lifting object-centric images and videos into 3D and 4D through rendering and score distillation objectives. A question then remains: what about lifting complete multi-object…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Wen-Hsuan Chu , Lei Ke , Katerina Fragkiadaki

We introduce a new benchmark, TAPVid-3D, for evaluating the task of long-range Tracking Any Point in 3D (TAP-3D). While point tracking in two dimensions (TAP) has many benchmarks measuring performance on real-world videos, such as…

Computer Vision and Pattern Recognition · Computer Science 2024-08-28 Skanda Koppula , Ignacio Rocco , Yi Yang , Joe Heyward , João Carreira , Andrew Zisserman , Gabriel Brostow , Carl Doersch

The last several years have seen significant progress in using depth cameras for tracking articulated objects such as human bodies, hands, and robotic manipulators. Most approaches focus on tracking skeletal parameters of a fixed shape…

Computer Vision and Pattern Recognition · Computer Science 2017-11-23 Aaron Walsman , Weilin Wan , Tanner Schmidt , Dieter Fox

Egocentric videos provide valuable insights into human interactions with the physical world, which has sparked growing interest in the computer vision and robotics communities. A critical challenge in fully understanding the geometry and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Chengbo Yuan , Geng Chen , Li Yi , Yang Gao