English
Related papers

Related papers: SS4D: Native 4D Generative Model via Structured Sp…

200 papers

We present a video generation model that accurately reproduces object motion, changes in camera viewpoint, and new content that arises over time. Existing video generation methods often fail to produce new content as a function of time…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Tim Brooks , Janne Hellsten , Miika Aittala , Ting-Chun Wang , Timo Aila , Jaakko Lehtinen , Ming-Yu Liu , Alexei A. Efros , Tero Karras

Recent advances in diffusion models have revolutionized 2D and 3D content creation, yet generating photorealistic dynamic 4D scenes remains a significant challenge. Existing dynamic 4D generation methods typically rely on distilling…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Vinayak Gupta , Yunze Man , Yu-Xiong Wang

Reconstructing dynamic 3D scenes from 2D images and generating diverse views over time is challenging due to scene complexity and temporal dynamics. Despite advancements in neural implicit models, limitations persist: (i) Inadequate Scene…

Computer Vision and Pattern Recognition · Computer Science 2024-02-23 Zeyu Yang , Hongye Yang , Zijie Pan , Li Zhang

Modeling dynamic 3D scenes is challenging due to their high-dimensional nature, which requires aggregating information from multiple views to reconstruct time-evolving 3D geometry and motion. We present a novel multi-video 4D Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yonghan Lee , Tsung-Wei Huang , Shiv Gehlot , Jaehoon Choi , Guan-Ming Su , Dinesh Manocha

Immersive applications call for synthesizing spatiotemporal 4D content from casual videos without costly 3D supervision. Existing video-to-4D methods typically rely on manually annotated camera poses, which are labor-intensive and brittle…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Dongyue Lu , Ao Liang , Tianxin Huang , Xiao Fu , Yuyang Zhao , Baorui Ma , Liang Pan , Wei Yin , Lingdong Kong , Wei Tsang Ooi , Ziwei Liu

We present V2M4, a novel 4D reconstruction method that directly generates a usable 4D mesh animation asset from a single monocular video. Unlike existing approaches that rely on priors from multi-view image and video generation models, our…

Graphics · Computer Science 2025-07-30 Jianqi Chen , Biao Zhang , Xiangjun Tang , Peter Wonka

Generative adversarial models (GANs) continue to produce advances in terms of the visual quality of still images, as well as the learning of temporal correlations. However, few works manage to combine these two interesting capabilities for…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Gereon Fox , Ayush Tewari , Mohamed Elgharib , Christian Theobalt

Recently, the generation of dynamic 3D objects from a video has shown impressive results. Existing methods directly optimize Gaussians using whole information in frames. However, when dynamic regions are interwoven with static regions…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Liying Yang , Chen Liu , Zhenwei Zhu , Ajian Liu , Hui Ma , Jian Nong , Yanyan Liang

We present a latent diffusion model over 3D scenes, that can be trained using only 2D image data. To achieve this, we first design an autoencoder that maps multi-view images to 3D Gaussian splats, and simultaneously builds a compressed…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Paul Henderson , Melonie de Almeida , Daniela Ivanova , Titas Anciukevičius

We present Stable Video 3D (SV3D) -- a latent video diffusion model for high-resolution, image-to-multi-view generation of orbital videos around a 3D object. Recent work on 3D generation propose techniques to adapt 2D generative models for…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Vikram Voleti , Chun-Han Yao , Mark Boss , Adam Letts , David Pankratz , Dmitry Tochilkin , Christian Laforte , Robin Rombach , Varun Jampani

We present 4DNeX, the first feed-forward framework for generating 4D (i.e., dynamic 3D) scene representations from a single image. In contrast to existing methods that rely on computationally intensive optimization or require multi-frame…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Zhaoxi Chen , Tianqi Liu , Long Zhuo , Jiawei Ren , Zeng Tao , He Zhu , Fangzhou Hong , Liang Pan , Ziwei Liu

3D Gaussian Splatting (3DGS) has garnered significant attention due to its superior scene representation fidelity and real-time rendering performance, especially for dynamic 3D scene reconstruction (\textit{i.e.}, 4D reconstruction).…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Henan Wang , Hanxin Zhu , Xinliang Gong , Tianyu He , Xin Li , Zhibo Chen

Spatio-temporal scene graphs provide a principled representation for modeling evolving object interactions, yet existing methods remain fundamentally frame-centric: they reason only about currently visible objects, discard entities upon…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Rohith Peddi , Saurabh , Shravan Shanmugam , Likhitha Pallapothula , Yu Xiang , Parag Singla , Vibhav Gogate

We present Motion 3-to-4, a feed-forward framework for synthesising high-quality 4D dynamic objects from a single monocular video and an optional 3D reference mesh. While recent advances have significantly improved 2D, video, and 3D content…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Hongyuan Chen , Xingyu Chen , Youjia Zhang , Zexiang Xu , Anpei Chen

We present LidarDM, a novel LiDAR generative model capable of producing realistic, layout-aware, physically plausible, and temporally coherent LiDAR videos. LidarDM stands out with two unprecedented capabilities in LiDAR generative…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Vlas Zyrianov , Henry Che , Zhijian Liu , Shenlong Wang

Recent advances in diffusion-based generative models have established a new paradigm for image and video relighting. However, extending these capabilities to 4D relighting remains challenging, due primarily to the scarcity of paired 4D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Zhenghuang Wu , Kang Chen , Zeyu Zhang , Hao Tang

We pose a new problem, In-2-4D, for generative 4D (i.e., 3D + motion) inbetweening to interpolate two single-view images. In contrast to video/4D generation from only text or a single image, our interpolative task can leverage more precise…

Graphics · Computer Science 2025-09-30 Sauradip Nag , Daniel Cohen-Or , Hao Zhang , Ali Mahdavi-Amiri

Given a monocular video, the goal of video re-rendering is to generate views of the scene from a novel camera trajectory. Existing methods face two distinct challenges. Geometrically unconditioned models lack spatial awareness, leading to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Mingyang Xie , Numair Khan , Tianfu Wang , Naina Dhingra , Seonghyeon Nam , Haitao Yang , Zhuo Hui , Christopher Metzler , Andrea Vedaldi , Hamed Pirsiavash , Lei Luo

Understanding dynamic scenes from casual videos is critical for scalable robot learning, yet four-dimensional (4D) reconstruction under strictly monocular settings remains highly ill-posed. To address this challenge, our key insight is that…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Can Li , Jie Gu , Jingmin Chen , Fangzhou Qiu , Lei Sun

Text-to-4D generation has recently been demonstrated viable by integrating a 2D image diffusion model with a video diffusion model. However, existing models tend to produce results with inconsistent motions and geometric structures over…

Graphics · Computer Science 2024-08-19 Ce Chen , Shaoli Huang , Xuelin Chen , Guangyi Chen , Xiaoguang Han , Kun Zhang , Mingming Gong