English
Related papers

Related papers: Any4D: Unified Feed-Forward Metric 4D Reconstructi…

200 papers

Recent advancements in diffusion models for 2D and 3D content creation have sparked a surge of interest in generating 4D content. However, the scarcity of 3D scene datasets constrains current methodologies to primarily object-centric…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Dejia Xu , Hanwen Liang , Neel P. Bhatt , Hezhen Hu , Hanxue Liang , Konstantinos N. Plataniotis , Zhangyang Wang

Models for egocentric 3D and 4D reconstruction, including few-shot interpolation and extrapolation settings, can benefit from having images from exocentric viewpoints as supervision signals. No existing dataset provides the necessary…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Marius Kästingschäfer , Théo Gieruc , Sebastian Bernhard , Dylan Campbell , Eldar Insafutdinov , Eyvaz Najafli , Thomas Brox

In this paper, we present Consistent4D, a novel approach for generating 4D dynamic objects from uncalibrated monocular videos. Uniquely, we cast the 360-degree dynamic object reconstruction as a 4D generation problem, eliminating the need…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Yanqin Jiang , Li Zhang , Jin Gao , Weimin Hu , Yao Yao

A long-standing goal in scene understanding is to obtain interpretable and editable representations that can be directly constructed from a raw monocular RGB-D video, without requiring specialized hardware setup or priors. The problem is…

Computer Vision and Pattern Recognition · Computer Science 2023-06-22 Yu-Shiang Wong , Niloy J. Mitra

In the last year, universal monocular metric depth estimation (universal MMDE) has gained considerable attention, serving as the foundation model for various multimedia tasks, such as video and image editing. Nonetheless, current approaches…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Yihao Liu , Feng Xue , Anlong Ming , Mingshuai Zhao , Huadong Ma , Nicu Sebe

Generative video models have significantly advanced the photorealistic synthesis of adverse weather for autonomous driving; however, they consistently demand massive datasets to learn rare weather scenarios. While 3D-aware editing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Tianyu Liu , Weitao Xiong , Kunming Luo , Manyuan Zhang , Peng Li , Yuan Liu , Ping Tan

We present a novel framework for dynamic radiance field prediction given monocular video streams. Unlike previous methods that primarily focus on predicting future frames, our method goes a step further by generating explicit 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Di Qi , Tong Yang , Beining Wang , Xiangyu Zhang , Wenqiang Zhang

Current compositional image-to-3D scene generation approaches construct 3D scenes by time-consuming iterative layout optimization or inflexible joint object-layout generation. Moreover, most methods rely on limited field-of-view perspective…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Zidian Qiu , Ancong Wu

Feedforward Gaussian Splatting has recently emerged as an efficient paradigm for 4D reconstruction in autonomous driving. However, in unstructured off-road scenes, its performance degrades due to high-frequency geometry, ego-motion jitter,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Shuo Wang , Jilin Mei , Fuyang Liu , Wenfei Guan , Fanjie Kong , Zhihua Zhao , Shuai Wang , Chen Min , Yu Hu

Dynamic driving scene reconstruction is critical for autonomous driving simulation and closed-loop learning. While recent feed-forward methods have shown promise for 3D reconstruction, they struggle with long-range driving sequences due to…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Kaiyuan Tan , Yingying Shen , Mingfei Tu , Haohui Zhu , Bing Wang , Guang Chen , Hangjun Ye , Haiyang Sun

Existing techniques for dynamic scene reconstruction from multiple wide-baseline cameras primarily focus on reconstruction in controlled environments, with fixed calibrated cameras and strong prior constraints. This paper introduces a…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Armin Mustafa , Marco Volino , Hansung Kim , Jean-Yves Guillemaut , Adrian Hilton

We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panoptic Studio). Such multi-view setups are prohibitively…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zihan Wang , Jeff Tan , Tarasha Khurana , Neehar Peri , Deva Ramanan

Recent advances in diffusion-based generative models have established a new paradigm for image and video relighting. However, extending these capabilities to 4D relighting remains challenging, due primarily to the scarcity of paired 4D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Zhenghuang Wu , Kang Chen , Zeyu Zhang , Hao Tang

Monocular depth estimation aims to recover the depth information of 3D scenes from 2D images. Recent work has made significant progress, but its reliance on large-scale datasets and complex decoders has limited its efficiency and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Zeyu Ren , Zeyu Zhang , Wukai Li , Qingxiang Liu , Hao Tang

In the realm of robot-assisted minimally invasive surgery, dynamic scene reconstruction can significantly enhance downstream tasks and improve surgical outcomes. Neural Radiance Fields (NeRF)-based methods have recently risen to prominence…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Yiming Huang , Beilei Cui , Long Bai , Ziqi Guo , Mengya Xu , Mobarakol Islam , Hongliang Ren

This paper targets high-fidelity and real-time view synthesis of dynamic 3D scenes at 4K resolution. Recently, some methods on dynamic view synthesis have shown impressive rendering quality. However, their speed is still limited when…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhen Xu , Sida Peng , Haotong Lin , Guangzhao He , Jiaming Sun , Yujun Shen , Hujun Bao , Xiaowei Zhou

High-fidelity rendering of dynamic humans from monocular videos typically degrades catastrophically under occlusions. Existing solutions incorporate external priors-either hallucinating missing content via generative models, which induces…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Weiquan Wang , Feifei Shao , Lin Li , Zhen Wang , Jun Xiao , Long Chen

Foundation models pre-trained on massive unlabeled datasets have revolutionized natural language and computer vision, exhibiting remarkable generalization capabilities, thus highlighting the importance of pre-training. Yet, efforts in…

Robotics · Computer Science 2025-05-20 Dantong Niu , Yuvan Sharma , Haoru Xue , Giscard Biamby , Junyi Zhang , Ziteng Ji , Trevor Darrell , Roei Herzig

Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images. However, existing active…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Tianling Xu , Shengzhe Gan , Leslie Gu , Yuelei Li , Fangneng Zhan , Hanspeter Pfister

In this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Chi Zhang , Wei Yin , Gang Yu , Zhibin Wang , Tao Chen , Bin Fu , Joey Tianyi Zhou , Chunhua Shen