中文
相关论文

相关论文: 4DGT: Learning a 4D Gaussian Transformer Using Rea…

200 篇论文

We present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets to enable novel view synthesis at any specified camera…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Rundi Wu , Ruiqi Gao , Ben Poole , Alex Trevithick , Changxi Zheng , Jonathan T. Barron , Aleksander Holynski

Reconstructing deformable tissues from endoscopic videos is essential in many downstream surgical applications. However, existing methods suffer from slow rendering speed, greatly limiting their practical use. In this paper, we introduce…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Yifan Liu , Chenxin Li , Chen Yang , Yixuan Yuan

Reconstructing dynamic 3D scenes from monocular video has broad applications in AR/VR, robotics, and autonomous navigation, but often fails due to severe motion blur caused by camera and object motion. Existing methods commonly follow a…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zhijing Wu , Longguang Wang

Recently, Gaussian Splatting methods have emerged as a desirable substitute for prior Radiance Field methods for novel-view synthesis of scenes captured with multi-view images or videos. In this work, we propose a novel extension to 4D…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Karly Hou , Wanhua Li , Hanspeter Pfister

We present Orientation-anchored Gaussian Splatting (OriGS), a novel framework for high-quality 4D reconstruction from casually captured monocular videos. While recent advances extend 3D Gaussian Splatting to dynamic scenes via various…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Junyi Wu , Jiachen Tao , Haoxuan Wang , Gaowen Liu , Ramana Rao Kompella , Yan Yan

We present L4GM, the first 4D Large Reconstruction Model that produces animated objects from a single-view video input -- in a single feed-forward pass that takes only a second. Key to our success is a novel dataset of multiview videos…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Jiawei Ren , Kevin Xie , Ashkan Mirzaei , Hanxue Liang , Xiaohui Zeng , Karsten Kreis , Ziwei Liu , Antonio Torralba , Sanja Fidler , Seung Wook Kim , Huan Ling

Reconstructing dynamic humans together with static scenes from monocular videos remains difficult, especially under fast motion, where RGB frames suffer from motion blur. Event cameras exhibit distinct advantages, e.g., microsecond temporal…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Xiaoting Yin , Hao Shi , Kailun Yang , Jiajun Zhai , Shangwei Guo , Lin Wang , Kaiwei Wang

Common computer vision systems typically assume ideal pinhole cameras but fail when facing real-world camera effects such as fisheye distortion and rolling shutter, mainly due to the lack of learning from training data with camera effects.…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Yi-Ruei Liu , You-Zhe Xie , Yu-Hsiang Hsu , I-Sheng Fang , Yu-Lun Liu , Jun-Cheng Chen

Recent advancements in 2D/3D generative techniques have facilitated the generation of dynamic 3D objects from monocular videos. Previous methods mainly rely on the implicit neural radiance fields (NeRF) or explicit Gaussian Splatting as the…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Zhiqi Li , Yiming Chen , Peidong Liu

Recent advances in 2D/3D generative models enable the generation of dynamic 3D objects from a single-view video. Existing approaches utilize score distillation sampling to form the dynamic scene as dynamic NeRF or dense 3D Gaussians.…

计算机视觉与模式识别 · 计算机科学 2024-08-15 Zijie Wu , Chaohui Yu , Yanqin Jiang , Chenjie Cao , Fan Wang , Xiang Bai

Egocentric video is crucial for next-generation 4D scene reconstruction, with applications in AR/VR and embodied AI. However, reconstructing dynamic first-person scenes is challenging due to complex ego-motion, occlusions, and hand-object…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Tingxi Chen , Zhengxue Cheng , Houqiang Zhong , Su Wang , Rong Xie , Li Song

Rendering dynamic scenes from monocular videos is a crucial yet challenging task. The recent deformable Gaussian Splatting has emerged as a robust solution to represent real-world dynamic scenes. However, it often leads to heavily redundant…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Hanyang Kong , Xingyi Yang , Xinchao Wang

We consider the problem of efficiently representing casually captured monocular videos in a spatially- and temporally-coherent manner. While existing approaches predominantly rely on 2D/2.5D techniques treating videos as collections of…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Qiuhong Shen , Xuanyu Yi , Mingbao Lin , Hanwang Zhang , Shuicheng Yan , Xinchao Wang

Remarkable advances in recent 2D image and 3D shape generation have induced a significant focus on dynamic 4D content generation. However, previous 4D generation methods commonly struggle to maintain spatial-temporal consistency and adapt…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Mengmeng Liu , Jiuming Liu , Yunpeng Zhang , Jiangtao Li , Michael Ying Yang , Francesco Nex , Hao Cheng

The accurate reconstruction of dynamic street scenes is critical for applications in autonomous driving, augmented reality, and virtual reality. Traditional methods relying on dense point clouds and triangular meshes struggle with moving…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Peizhen Zheng , Dongjing Jiang , Qingchong Jiao , Redouane EL Bouchtaoui , Flynnwell Jianfei Zhang

Existing dynamic scene reconstruction methods based on Gaussian Splatting enable real-time rendering and generate realistic images. However, adjusting the camera's focal length or the distance between Gaussian primitives and the camera to…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Zilong Chen , Huan-ang Gao , Delin Qu , Haohan Chi , Hao Tang , Kai Zhang , Hao Zhao

Reconstructing dynamic scenes with large-scale and complex motions remains a significant challenge. Recent techniques like Neural Radiance Fields and 3D Gaussian Splatting (3DGS) have shown promise but still struggle with scenes involving…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Qiankun Gao , Yanmin Wu , Chengxiang Wen , Jiarui Meng , Luyang Tang , Jie Chen , Ronggang Wang , Jian Zhang

4D generation has made remarkable progress in synthesizing dynamic 3D objects from input text, images, or videos. However, existing methods often represent motion as an implicit deformation field, which limits direct control and…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Lifan Wu , Ruijie Zhu , Yubo Ai , Tianzhu Zhang

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes,…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Zeyi Liu , Shuang Li , Eric Cousineau , Siyuan Feng , Benjamin Burchfiel , Shuran Song

Dynamic 3D scene representation and novel view synthesis are crucial for enabling immersive experiences required by AR/VR and metaverse applications. It is a challenging task due to the complexity of unconstrained real-world scenes and…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Zeyu Yang , Zijie Pan , Xiatian Zhu , Li Zhang , Jianfeng Feng , Yu-Gang Jiang , Philip H. S. Torr