English
Related papers

Related papers: Any4D: Unified Feed-Forward Metric 4D Reconstructi…

200 papers

A foundational humanoid motion tracker is expected to be able to track diverse, highly dynamic, and contact-rich motions. More importantly, it needs to operate stably in real-world scenarios against various dynamics disturbances, including…

Optical flow estimation is a crucial subfield of computer vision, serving as a foundation for video tasks. However, the real-world robustness is limited by animated synthetic datasets for training. This introduces domain gaps when applied…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Yingping Liang , Ying Fu , Yutao Hu , Wenqi Shao , Jiaming Liu , Debing Zhang

Novel view synthesis from monocular videos of dynamic scenes with unknown camera poses remains a fundamental challenge in computer vision and graphics. While recent advances in 3D representations such as Neural Radiance Fields (NeRF) and 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Mengqi Guo , Bo Xu , Yanyan Li , Gim Hee Lee

We propose a novel method to accurately reconstruct a set of images representing a single scene from few linear multi-view measurements. Each observed image is modeled as the sum of a background image and a foreground one. The background…

Computer Vision and Pattern Recognition · Computer Science 2013-09-19 Gilles Puy , Pierre Vandergheynst

Recent feed-forward networks have achieved remarkable progress in sparse-view 3D reconstruction by predicting dense point maps directly from RGB images. However, they often suffer from geometric inconsistencies and limited fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yutong Chen , Yiming Wang , Xucong Zhang , Sergey Prokudin , Siyu Tang

Precise camera control for reshooting dynamic videos is bottlenecked by the severe scarcity of paired multi-view data for non-rigid scenes. We overcome this limitation with a highly scalable self-supervised framework capable of leveraging…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Avinash Paliwal , Adithya Iyer , Shivin Yadav , Muhammad Ali Afridi , Midhun Harikumar

Generating interactive and dynamic 4D scenes from a single static image remains a core challenge. Most existing generate-then-reconstruct and reconstruct-then-generate methods decouple geometry from motion, causing spatiotemporal…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yanran Zhang , Ziyi Wang , Wenzhao Zheng , Zheng Zhu , Jie Zhou , Jiwen Lu

3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise and material sensitivity, while existing reconstruction models…

Robotics · Computer Science 2026-05-05 Sizhe Yang , Linning Xu , Hao Li , Juncheng Mu , Jia Zeng , Dahua Lin , Jiangmiao Pang

Live-streaming Novel View Synthesis (NVS) from unposed multi-view video remains an open challenge in a wide range of applications. Existing methods for dynamic scene representation typically require ground-truth camera parameters and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Pedro Quesado , Erkut Akdag , Yasaman Kashefbahrami , Willem Menu , Egor Bondarev

We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Jianyuan Wang , Minghao Chen , Nikita Karaev , Andrea Vedaldi , Christian Rupprecht , David Novotny

Feed-forward paradigms for 3D reconstruction have become a focus of recent research, which learn implicit, fixed view transformations to generate a single scene representation. However, their application to complex driving scenes reveals…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Haochen Yu , Qiankun Liu , Hongyuan Liu , Jianfei Jiang , Juntao Lyu , Jiansheng Chen , Huimin Ma

Reconstructing a complete 3D head from a single portrait remains challenging because existing methods still face a sharp quality-speed trade-off: high-fidelity pipelines often rely on multi-stage processing and per-subject optimization,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Yujie Gao , Yao Xiao , Xiangnan Zhu , Ya Li , Yiyi Zhang , Liqing Zhang , Jianfu Zhang

3D scene reconstruction is a long-standing vision task. Existing approaches can be categorized into geometry-based and learning-based methods. The former leverages multi-view geometry but can face catastrophic failures due to the reliance…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Guangkai Xu , Wei Yin , Hao Chen , Chunhua Shen , Kai Cheng , Feng Zhao

Immersive applications call for synthesizing spatiotemporal 4D content from casual videos without costly 3D supervision. Existing video-to-4D methods typically rely on manually annotated camera poses, which are labor-intensive and brittle…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Dongyue Lu , Ao Liang , Tianxin Huang , Xiao Fu , Yuyang Zhao , Baorui Ma , Liang Pan , Wei Yin , Lingdong Kong , Wei Tsang Ooi , Ziwei Liu

Reconstructing the 3D shape of a deformable environment from the information captured by a moving depth camera is highly relevant to surgery. The underlying challenge is the fact that simultaneously estimating camera motion and tissue…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Guido Caccianiga , Julian Nubert , Cesar Cadena , Marco Hutter , Katherine J. Kuchenbecker

Recent advancements in static feed-forward scene reconstruction have demonstrated significant progress in high-quality novel view synthesis. However, these models often struggle with generalizability across diverse environments and fail to…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Hanxue Liang , Jiawei Ren , Ashkan Mirzaei , Antonio Torralba , Ziwei Liu , Igor Gilitschenski , Sanja Fidler , Cengiz Oztireli , Huan Ling , Zan Gojcic , Jiahui Huang

Planning and acting in 3D environments is a fundamental capability for robotic manipulation in the real world. Although prior work has explored predictive flow planners to guide 3D manipulation, existing approaches often rely on modular…

Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-modal translation methods treat each modality pair as an…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Haoyang Chen , Jing Zhang , Hebaixu Wang , Shiqin Wang , Pohsun Huang , Jiayuan Li , Haonan Guo , Di Wang , Zheng Wang , Bo Du

We present a method for depth estimation with monocular images, which can predict high-quality depth on diverse scenes up to an affine transformation, thus preserving accurate shapes of a scene. Previous methods that predict metric depth…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Wei Yin , Xinlong Wang , Chunhua Shen , Yifan Liu , Zhi Tian , Songcen Xu , Changming Sun , Dou Renyin

Generating high-quality 3D content from text, single images, or sparse view images remains a challenging task with broad applications. Existing methods typically employ multi-view diffusion models to synthesize multi-view images, followed…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Junlin Han , Jianyuan Wang , Andrea Vedaldi , Philip Torr , Filippos Kokkinos