English
Related papers

Related papers: 4DGT: Learning a 4D Gaussian Transformer Using Rea…

200 papers

We present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets to enable novel view synthesis at any specified camera…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Rundi Wu , Ruiqi Gao , Ben Poole , Alex Trevithick , Changxi Zheng , Jonathan T. Barron , Aleksander Holynski

Reconstructing deformable tissues from endoscopic videos is essential in many downstream surgical applications. However, existing methods suffer from slow rendering speed, greatly limiting their practical use. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-02-14 Yifan Liu , Chenxin Li , Chen Yang , Yixuan Yuan

Reconstructing dynamic 3D scenes from monocular video has broad applications in AR/VR, robotics, and autonomous navigation, but often fails due to severe motion blur caused by camera and object motion. Existing methods commonly follow a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zhijing Wu , Longguang Wang

Recently, Gaussian Splatting methods have emerged as a desirable substitute for prior Radiance Field methods for novel-view synthesis of scenes captured with multi-view images or videos. In this work, we propose a novel extension to 4D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Karly Hou , Wanhua Li , Hanspeter Pfister

We present Orientation-anchored Gaussian Splatting (OriGS), a novel framework for high-quality 4D reconstruction from casually captured monocular videos. While recent advances extend 3D Gaussian Splatting to dynamic scenes via various…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Junyi Wu , Jiachen Tao , Haoxuan Wang , Gaowen Liu , Ramana Rao Kompella , Yan Yan

We present L4GM, the first 4D Large Reconstruction Model that produces animated objects from a single-view video input -- in a single feed-forward pass that takes only a second. Key to our success is a novel dataset of multiview videos…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Jiawei Ren , Kevin Xie , Ashkan Mirzaei , Hanxue Liang , Xiaohui Zeng , Karsten Kreis , Ziwei Liu , Antonio Torralba , Sanja Fidler , Seung Wook Kim , Huan Ling

Reconstructing dynamic humans together with static scenes from monocular videos remains difficult, especially under fast motion, where RGB frames suffer from motion blur. Event cameras exhibit distinct advantages, e.g., microsecond temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Xiaoting Yin , Hao Shi , Kailun Yang , Jiajun Zhai , Shangwei Guo , Lin Wang , Kaiwei Wang

Common computer vision systems typically assume ideal pinhole cameras but fail when facing real-world camera effects such as fisheye distortion and rolling shutter, mainly due to the lack of learning from training data with camera effects.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Yi-Ruei Liu , You-Zhe Xie , Yu-Hsiang Hsu , I-Sheng Fang , Yu-Lun Liu , Jun-Cheng Chen

Recent advancements in 2D/3D generative techniques have facilitated the generation of dynamic 3D objects from monocular videos. Previous methods mainly rely on the implicit neural radiance fields (NeRF) or explicit Gaussian Splatting as the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Zhiqi Li , Yiming Chen , Peidong Liu

Recent advances in 2D/3D generative models enable the generation of dynamic 3D objects from a single-view video. Existing approaches utilize score distillation sampling to form the dynamic scene as dynamic NeRF or dense 3D Gaussians.…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Zijie Wu , Chaohui Yu , Yanqin Jiang , Chenjie Cao , Fan Wang , Xiang Bai

Egocentric video is crucial for next-generation 4D scene reconstruction, with applications in AR/VR and embodied AI. However, reconstructing dynamic first-person scenes is challenging due to complex ego-motion, occlusions, and hand-object…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Tingxi Chen , Zhengxue Cheng , Houqiang Zhong , Su Wang , Rong Xie , Li Song

Rendering dynamic scenes from monocular videos is a crucial yet challenging task. The recent deformable Gaussian Splatting has emerged as a robust solution to represent real-world dynamic scenes. However, it often leads to heavily redundant…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Hanyang Kong , Xingyi Yang , Xinchao Wang

We consider the problem of efficiently representing casually captured monocular videos in a spatially- and temporally-coherent manner. While existing approaches predominantly rely on 2D/2.5D techniques treating videos as collections of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Qiuhong Shen , Xuanyu Yi , Mingbao Lin , Hanwang Zhang , Shuicheng Yan , Xinchao Wang

Remarkable advances in recent 2D image and 3D shape generation have induced a significant focus on dynamic 4D content generation. However, previous 4D generation methods commonly struggle to maintain spatial-temporal consistency and adapt…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Mengmeng Liu , Jiuming Liu , Yunpeng Zhang , Jiangtao Li , Michael Ying Yang , Francesco Nex , Hao Cheng

The accurate reconstruction of dynamic street scenes is critical for applications in autonomous driving, augmented reality, and virtual reality. Traditional methods relying on dense point clouds and triangular meshes struggle with moving…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Peizhen Zheng , Dongjing Jiang , Qingchong Jiao , Redouane EL Bouchtaoui , Flynnwell Jianfei Zhang

Existing dynamic scene reconstruction methods based on Gaussian Splatting enable real-time rendering and generate realistic images. However, adjusting the camera's focal length or the distance between Gaussian primitives and the camera to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Zilong Chen , Huan-ang Gao , Delin Qu , Haohan Chi , Hao Tang , Kai Zhang , Hao Zhao

Reconstructing dynamic scenes with large-scale and complex motions remains a significant challenge. Recent techniques like Neural Radiance Fields and 3D Gaussian Splatting (3DGS) have shown promise but still struggle with scenes involving…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Qiankun Gao , Yanmin Wu , Chengxiang Wen , Jiarui Meng , Luyang Tang , Jie Chen , Ronggang Wang , Jian Zhang

4D generation has made remarkable progress in synthesizing dynamic 3D objects from input text, images, or videos. However, existing methods often represent motion as an implicit deformation field, which limits direct control and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Lifan Wu , Ruijie Zhu , Yubo Ai , Tianzhu Zhang

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Zeyi Liu , Shuang Li , Eric Cousineau , Siyuan Feng , Benjamin Burchfiel , Shuran Song

Dynamic 3D scene representation and novel view synthesis are crucial for enabling immersive experiences required by AR/VR and metaverse applications. It is a challenging task due to the complexity of unconstrained real-world scenes and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Zeyu Yang , Zijie Pan , Xiatian Zhu , Li Zhang , Jianfeng Feng , Yu-Gang Jiang , Philip H. S. Torr