English
Related papers

Related papers: Generative 4D Scene Gaussian Splatting with Object…

200 papers

We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene-level 3D generation. We repurpose the VGGT reconstruction model to produce geometric latents by training an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Jiaxin Huang , Yuanbo Yang , Bangbang Yang , Lin Ma , Yuewen Ma , Yiyi Liao

4D head capture aims to generate dynamic topological meshes and corresponding texture maps from videos, which is widely utilized in movies and games for its ability to simulate facial muscle movements and recover dynamic textures in…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Xuanchen Li , Yuhao Cheng , Xingyu Ren , Haozhe Jia , Di Xu , Wenhan Zhu , Yichao Yan

Given a video and a set of input object masks, an omnimatte method aims to decompose the video into semantically meaningful layers containing individual objects along with their associated effects, such as shadows and reflections. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yao-Chih Lee , Erika Lu , Sarah Rumbley , Michal Geyer , Jia-Bin Huang , Tali Dekel , Forrester Cole

Scene reconstruction and novel-view synthesis for large, complex, multi-story, indoor scenes is a challenging and time-consuming task. Prior methods have utilized drones for data capture and radiance fields for scene reconstruction, both of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Yuanbo Chen , Chengyu Zhang , Jason Wang , Xuefan Gao , Avideh Zakhor

Multi-view video reconstruction plays a vital role in computer vision, enabling applications in film production, virtual reality, and motion analysis. While recent advances such as 4D Gaussian Splatting (4DGS) have demonstrated impressive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Zhixin Xu , Hengyu Zhou , Yuan Liu , Wenhan Xue , Hao Pan , Wenping Wang , Bin Wang

Computer vision-based technologies significantly enhance surgical automation by advancing tool tracking, detection, and localization. However, Current data-driven approaches are data-voracious, requiring large, high-quality labeled image…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Tianle Zeng , Junlei Hu , Gerardo Loza Galindo , Sharib Ali , Duygu Sarikaya , Pietro Valdastri , Dominic Jones

The reconstruction of dynamic 3D scenes using 3D Gaussian Splatting has shown significant promise. A key challenge, however, remains in modeling realistic motion, as most methods fail to align the motion of Gaussians with real-world…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Junoh Lee , Junmyeong Lee , Yeon-Ji Song , Inhwan Bae , Jisu Shin , Hae-Gon Jeon , Jin-Hwa Kim

We present GP-4DGS, a novel framework that integrates Gaussian Processes (GPs) into 4D Gaussian Splatting (4DGS) for principled probabilistic modeling of dynamic scenes. While existing 4DGS methods focus on deterministic reconstruction,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Mijeong Kim , Jungtaek Kim , Bohyung Han

With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Xiaolu Hou , Mingcheng Li , Dingkang Yang , Jiawei Chen , Ziyun Qian , Xiao Zhao , Yue Jiang , Jinjie Wei , Qingyao Xu , Lihua Zhang

3D Gaussian Splatting (3DGS) has revolutionized 3D scene reconstruction, which effectively balances rendering quality, efficiency, and speed. However, existing 3DGS approaches usually generate plausible outputs and face significant…

Graphics · Computer Science 2025-07-22 Beizhen Zhao , Yifan Zhou , Sicheng Yu , Zijian Wang , Hao Wang

Novel view synthesis from unconstrained in-the-wild images remains a meaningful but challenging task. The photometric variation and transient occluders in those unconstrained images make it difficult to reconstruct the original scene…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Dongbin Zhang , Chuming Wang , Weitao Wang , Peihao Li , Minghan Qin , Haoqian Wang

We present MOSAIC-GS, a novel, fully explicit, and computationally efficient approach for high-fidelity dynamic scene reconstruction from monocular videos using Gaussian Splatting. Monocular reconstruction is inherently ill-posed due to the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Svitlana Morkva , Maximum Wilder-Smith , Michael Oechsle , Alessio Tonioni , Marco Hutter , Vaishakh Patil

We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic priors captured by large-scale pre-trained video models, Geo4D can be trained using only…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Zeren Jiang , Chuanxia Zheng , Iro Laina , Diane Larlus , Andrea Vedaldi

We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a…

Recent 3D reconstruction methods achieve impressive results with dense multi-view imagery but struggle when only a few views are available. Various approaches, including regularization techniques, semantic priors, and geometric constraints,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Yi-Chuan Huang , Hao-Jen Chien , Chin-Yang Lin , Ying-Huan Chen , Yu-Lun Liu

Reconstructing dynamic 3D scenes from sparse multi-view videos is highly ill-posed, often leading to geometric collapse, trajectory drift, and floating artifacts. Recent attempts introduce generative priors to hallucinate missing content,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Zhenlong Wu , Zihan Zheng , Xuanxuan Wang , Qianhe Wang , Hua Yang , Xiaoyun Zhang , Qiang Hu , Wenjun Zhang

We propose VideoRFSplat, a direct text-to-3D model leveraging a video generation model to generate realistic 3D Gaussian Splatting (3DGS) for unbounded real-world scenes. To generate diverse camera poses and unbounded spatial extent of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Hyojun Go , Byeongjun Park , Hyelin Nam , Byung-Hoon Kim , Hyungjin Chung , Changick Kim

We introduce a novel method for dynamic free-view synthesis of an ambient scenes from a monocular capture bringing a immersive quality to the viewing experience. Our method builds upon the recent advancements in 3D Gaussian Splatting (3DGS)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Meng-Li Shih , Jia-Bin Huang , Changil Kim , Rajvi Shah , Johannes Kopf , Chen Gao

High-dynamic scene reconstruction aims to represent static background with rigid spatial features and dynamic objects with deformed continuous spatiotemporal features. Typically, existing methods adopt unified representation model (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Hanyu Zhou , Haonan Wang , Haoyue Liu , Yuxing Duan , Luxin Yan , Gim Hee Lee

The increasing demand for virtual reality applications has highlighted the significance of crafting immersive 3D assets. We present a text-to-3D 360$^{\circ}$ scene generation pipeline that facilitates the creation of comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Shijie Zhou , Zhiwen Fan , Dejia Xu , Haoran Chang , Pradyumna Chari , Tejas Bharadwaj , Suya You , Zhangyang Wang , Achuta Kadambi