English
Related papers

Related papers: FreeOrbit4D: Training-Free Arbitrary Camera Redire…

200 papers

Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often struggle with fine-grained physical consistency, exhibiting physically implausible dynamics over time.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Haoran Lu , Shang Wu , Jianshu Zhang , Maojiang Su , Guo Ye , Chenwei Xu , Lie Lu , Pranav Maneriker , Fan Du , Manling Li , Zhaoran Wang , Han Liu

Reconstructing deformable surgical scenes from endoscopic videos is challenging and clinically important. Recent state-of-the-art methods based on implicit neural representations or 3D Gaussian splatting have made notable progress. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Jiwei Shan , Zeyu Cai , Cheng-Tai Hsieh , Yirui Li , Hao Liu , Lijun Han , Hesheng Wang , Shing Shin Cheng

Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Linyi Jin , Richard Tucker , Zhengqi Li , David Fouhey , Noah Snavely , Aleksander Holynski

Immersive applications call for synthesizing spatiotemporal 4D content from casual videos without costly 3D supervision. Existing video-to-4D methods typically rely on manually annotated camera poses, which are labor-intensive and brittle…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Dongyue Lu , Ao Liang , Tianxin Huang , Xiao Fu , Yuyang Zhao , Baorui Ma , Liang Pan , Wei Yin , Lingdong Kong , Wei Tsang Ooi , Ziwei Liu

Despite significant advances in modeling image priors via diffusion models, 3D-aware image editing remains challenging, in part because the object is only specified via a single image. To tackle this challenge, we propose 3D-Fixup, a new…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Yen-Chi Cheng , Krishna Kumar Singh , Jae Shin Yoon , Alex Schwing , Liangyan Gui , Matheus Gadelha , Paul Guerrero , Nanxuan Zhao

We present ReFlow, a unified framework for monocular dynamic scene reconstruction that learns 3D motion in a novel self-correction manner from raw video. Existing methods often suffer from incomplete scene initialization for dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Yanzhe Liang , Ruijie Zhu , Hanzhi Chang , Zhuoyuan Li , Jiahao Lu , Tianzhu Zhang

We present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets to enable novel view synthesis at any specified camera…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Rundi Wu , Ruiqi Gao , Ben Poole , Alex Trevithick , Changxi Zheng , Jonathan T. Barron , Aleksander Holynski

Recent progress in video diffusion models has spurred growing interest in camera-controlled novel-view video generation for dynamic scenes, aiming to provide creators with cinematic camera control capabilities in post-production. A key…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Min-Jung Kim , Jeongho Kim , Hoiyeong Jin , Junha Hyung , Jaegul Choo

Dynamic view synthesis has seen significant advances, yet reconstructing scenes from uncalibrated, casual video remains challenging due to slow optimization and complex parameter estimation. In this work, we present Instant4D, a monocular…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Zhanpeng Luo , Haoxi Ran , Li Lu

Video diffusion models perform well in short-video synthesis, but their training-free extension to long videos often suffers from content drift, temporal inconsistency, and over-smoothed dynamics. Existing methods improve temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Fangda Chen , Shanshan Zhao , Longrong Yang , Chuanfu Xu , Zhigang Luo , Long Lan

Recent advancements in 2D and multimodal models have achieved remarkable success by leveraging large-scale training on extensive datasets. However, extending these achievements to enable free-form interactions and high-level semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Shijie Zhou , Hui Ren , Yijia Weng , Shuwang Zhang , Zhen Wang , Dejia Xu , Zhiwen Fan , Suya You , Zhangyang Wang , Leonidas Guibas , Achuta Kadambi

This paper presents a unified approach to understanding dynamic scenes from casual videos. Large pretrained vision foundation models, such as vision-language, video depth prediction, motion tracking, and segmentation models, offer promising…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 David Yifan Yao , Albert J. Zhai , Shenlong Wang

In this paper, we propose NeoVerse, a versatile 4D world model that is capable of 4D reconstruction, novel-trajectory video generation, and rich downstream applications. We first identify a common limitation of scalability in current 4D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yuxue Yang , Lue Fan , Ziqi Shi , Junran Peng , Feng Wang , Zhaoxiang Zhang

Diffusion models have recently emerged as powerful tools for camera simulation, enabling both geometric transformations and realistic optical effects. Among these, image-based bokeh rendering has shown promising results, but diffusion for…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Yang Yang , Siming Zheng , Qirui Yang , Jinwei Chen , Boxi Wu , Xiaofei He , Deng Cai , Bo Li , Peng-Tao Jiang

Despite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Wei Yin , Jianming Zhang , Oliver Wang , Simon Niklaus , Simon Chen , Yifan Liu , Chunhua Shen

Reconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide accurate 3D geometry, their performance drops markedly when…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yu Hu , Chong Cheng , Sicheng Yu , Xiaoyang Guo , Hao Wang

Recent progress in video generation has led to substantial improvements in visual fidelity, yet ensuring physically consistent motion remains a fundamental challenge. Intuitively, this limitation can be attributed to the fact that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Cong Wang , Hanxin Zhu , Xiao Tang , Jiayi Luo , Xin Jin , Long Chen , Zhibo Chen

Articulated 3D objects are central to many applications in robotics, AR/VR, and animation. Recent approaches to modeling such objects either rely on optimization-based reconstruction pipelines that require dense-view supervision or on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Chuhao Chen , Isabella Liu , Xinyue Wei , Hao Su , Minghua Liu

Reconstructing large-scale dynamic scenes from visual observations is a fundamental challenge in computer vision, with critical implications for robotics and autonomous systems. While recent differentiable rendering methods such as Neural…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Jingkang Wang , Henry Che , Yun Chen , Ze Yang , Lily Goli , Sivabalan Manivasagam , Raquel Urtasun

The ubiquity of monocular videos capturing daily hand-object interactions presents a valuable resource for embodied intelligence. While 3D hand reconstruction from in-the-wild videos has seen significant progress, reconstructing the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Yuantao Chen , Jiahao Chang , Chongjie Ye , Chaoran Zhang , Zhaojie Fang , Chenghong Li , Xiaoguang Han