English
Related papers

Related papers: 4RC: 4D Reconstruction via Conditional Querying An…

200 papers

Reconstructing and tracking dynamic 3D scenes remains a fundamental challenge in computer vision. Existing approaches often decouple geometry from motion: multi-view reconstruction methods assume static scenes, while dynamic tracking…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Shenhan Qian , Ganlin Zhang , Shangzhe Wu , Daniel Cremers

Reconstructing 4D dynamic scenes from casually captured monocular videos is valuable but highly challenging, as each timestamp is observed from a single viewpoint. We introduce Vivid4D, a novel approach that enhances 4D monocular video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Jiaxin Huang , Sheng Miao , BangBang Yang , Yuewen Ma , Yiyi Liao

We present Motion 3-to-4, a feed-forward framework for synthesising high-quality 4D dynamic objects from a single monocular video and an optional 3D reference mesh. While recent advances have significantly improved 2D, video, and 3D content…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Hongyuan Chen , Xingyu Chen , Youjia Zhang , Zexiang Xu , Anpei Chen

We present a method to reconstruct time-consistent human body models from monocular videos, focusing on extremely loose clothing or handheld object interactions. Prior work in human reconstruction is either limited to tight clothing with no…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Jeff Tan , Donglai Xiang , Shubham Tulsiani , Deva Ramanan , Gengshan Yang

Creating deformable 3D content has gained increasing attention with the rise of text-to-image and image-to-video generative models. While these models provide rich semantic priors for appearance, they struggle to capture the physical…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Jixuan He , Chieh Hubert Lin , Lu Qi , Ming-Hsuan Yang

Reconstructing deformable surgical scenes from endoscopic videos is challenging and clinically important. Recent state-of-the-art methods based on implicit neural representations or 3D Gaussian splatting have made notable progress. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Jiwei Shan , Zeyu Cai , Cheng-Tai Hsieh , Yirui Li , Hao Liu , Lijun Han , Hesheng Wang , Shing Shin Cheng

With the rapid development of 3D reconstruction technology, research in 4D reconstruction is also advancing, existing 4D reconstruction methods can generate high-quality 4D scenes. However, due to the challenges in acquiring multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Ling Yang , Kaixin Zhu , Juanxi Tian , Bohan Zeng , Mingbao Lin , Hongjuan Pei , Wentao Zhang , Shuicheng Yan

We introduce UniCon3R, a unified feed-forward framework for online human-scene 4D reconstruction from monocular video. Current feed-forward human-scene reconstruction methods suffer from artifacts, where bodies float above the ground or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Tanuj Sur , Shashank Tripathi , Nikos Athanasiou , Ha Linh Nguyen , Kai Xu , Michael J. Black , Angela Yao

We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal input,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Anagh Malik , Dorian Chan , Xiaoming Zhao , David B. Lindell , Oncel Tuzel , Jen-Hao Rick Chang

Dynamic view synthesis has seen significant advances, yet reconstructing scenes from uncalibrated, casual video remains challenging due to slow optimization and complex parameter estimation. In this work, we present Instant4D, a monocular…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Zhanpeng Luo , Haoxi Ran , Li Lu

We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a…

We present Tensor4D, an efficient yet effective approach to dynamic scene modeling. The key of our solution is an efficient 4D tensor decomposition method so that the dynamic scene can be directly represented as a 4D spatio-temporal tensor.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Ruizhi Shao , Zerong Zheng , Hanzhang Tu , Boning Liu , Hongwen Zhang , Yebin Liu

Given the high complexity of directly generating high-dimensional data such as 4D, we present 4DVD, a cascaded video diffusion model that generates 4D content in a decoupled manner. Unlike previous multi-view video methods that directly…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Shuzhou Yang , Xiaodong Cun , Xiaoyu Li , Yaowei Li , Jian Zhang

We present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets to enable novel view synthesis at any specified camera…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Rundi Wu , Ruiqi Gao , Ben Poole , Alex Trevithick , Changxi Zheng , Jonathan T. Barron , Aleksander Holynski

We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Tianqi Liu , Zihao Huang , Zhaoxi Chen , Guangcong Wang , Shoukang Hu , Liao Shen , Huiqiang Sun , Zhiguo Cao , Wei Li , Ziwei Liu

Scenes in the real world are often composed of several static and dynamic objects. Capturing their 4-dimensional structures, composition and spatio-temporal configuration in-the-wild, though extremely interesting, is equally hard.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Ahmet Berke Gokmen , Ajad Chhatkuli , Luc Van Gool , Danda Pani Paudel

The synthesis of spatiotemporally coherent 4D content presents fundamental challenges in computer vision, requiring simultaneous modeling of high-fidelity spatial representations and physically plausible temporal dynamics. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Xiaoyan Liu , Kangrui Li , Yuehao Song , Jiaxin Liu

We present a novel framework named NeuralRecon for real-time 3D scene reconstruction from a monocular video. Unlike previous methods that estimate single-view depth maps separately on each key-frame and fuse them later, we propose to…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Jiaming Sun , Yiming Xie , Linghao Chen , Xiaowei Zhou , Hujun Bao

We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic priors captured by large-scale pre-trained video models, Geo4D can be trained using only…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Zeren Jiang , Chuanxia Zheng , Iro Laina , Diane Larlus , Andrea Vedaldi

Reconstructing dense geometry for dynamic scenes from a monocular video is a critical yet challenging task. Recent memory-based methods enable efficient online reconstruction, but they fundamentally suffer from a Memory Demand Dilemma: The…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Xudong Cai , Shuo Wang , Peng Wang , Yongcai Wang , Zhaoxin Fan , Wanting Li , Tianbao Zhang , Jianrong Tao , Yeying Jin , Deying Li