English
Related papers

Related papers: HoloTime: Taming Video Diffusion Models for Panora…

200 papers

The task of video generation requires synthesizing visually realistic and temporally coherent video frames. Existing methods primarily use asynchronous auto-regressive models or synchronous diffusion models to address this challenge.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Mingzhen Sun , Weining Wang , Gen Li , Jiawei Liu , Jiahui Sun , Wanquan Feng , Shanshan Lao , SiYu Zhou , Qian He , Jing Liu

Reconstructing dynamic 3D scenes (i.e., 4D geometry) from monocular video is an important yet challenging problem. Conventional multiview geometry-based approaches often struggle with dynamic motion, whereas recent learning-based methods…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jinjie Mai , Wenxuan Zhu , Haozhe Liu , Bing Li , Cheng Zheng , Jürgen Schmidhuber , Bernard Ghanem

The video generation field has witnessed rapid improvements with the introduction of recent diffusion models. While these models have successfully enhanced appearance quality, they still face challenges in generating coherent and natural…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Yaosi Hu , Zhenzhong Chen , Chong Luo

We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a…

We present DiffIR2VR-Zero, a zero-shot framework that enables any pre-trained image restoration diffusion model to perform high-quality video restoration without additional training. While image diffusion models have shown remarkable…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Chang-Han Yeh , Hau-Shiang Shiu , Chin-Yang Lin , Zhixiang Wang , Chi-Wei Hsiao , Ting-Hsuan Chen , Yu-Lun Liu

Video outpainting is a challenging task that generates new video content by extending beyond the boundaries of an original input video, requiring both temporal and spatial consistency. Many state-of-the-art methods utilize latent diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Linhao Zhong , Fan Li , Yi Huang , Jianzhuang Liu , Renjing Pei , Fenglong Song

Text-driven 3D scene generation has seen significant advancements recently. However, most existing methods generate single-view images using generative models and then stitch them together in 3D space. This independent generation for each…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Wenrui Li , Fucheng Cai , Yapeng Mi , Zhe Yang , Wangmeng Zuo , Xingtao Wang , Xiaopeng Fan

Generating 3D scenes is a challenging open problem, which requires synthesizing plausible content that is fully consistent in 3D space. While recent methods such as neural radiance fields excel at view synthesis and 3D reconstruction, they…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Titas Anciukevičius , Fabian Manhardt , Federico Tombari , Paul Henderson

We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic priors captured by large-scale pre-trained video models, Geo4D can be trained using only…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Zeren Jiang , Chuanxia Zheng , Iro Laina , Diane Larlus , Andrea Vedaldi

Text-driven image and video diffusion models have recently achieved unprecedented generation realism. While diffusion models have been successfully applied for image editing, very few works have done so for video editing. We present the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-03 Eyal Molad , Eliahu Horwitz , Dani Valevski , Alex Rav Acha , Yossi Matias , Yael Pritch , Yaniv Leviathan , Yedid Hoshen

Diffusion models have emerged as the best approach for generative modeling of 2D images. Part of their success is due to the possibility of training them on millions if not billions of images with a stable learning objective. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Animesh Karnewar , Andrea Vedaldi , David Novotny , Niloy Mitra

Current video deblurring methods have limitations in recovering high-frequency information since the regression losses are conservative with high-frequency details. Since Diffusion Models (DMs) have strong capabilities in generating…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Chen Rao , Guangyuan Li , Zehua Lan , Jiakai Sun , Junsheng Luan , Wei Xing , Lei Zhao , Huaizhong Lin , Jianfeng Dong , Dalong Zhang

The increasing demand for augmented and virtual reality applications has highlighted the importance of crafting immersive 3D scenes from a simple single-view image. However, due to the partial priors provided by single-view input, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Tianyi Gong , Boyan Li , Yifei Zhong , Fangxin Wang

AI-generated content has attracted lots of attention recently, but photo-realistic video synthesis is still challenging. Although many attempts using GANs and autoregressive models have been made in this area, the visual quality and length…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Yingqing He , Tianyu Yang , Yong Zhang , Ying Shan , Qifeng Chen

Reconstructing dynamic 3D scenes from 2D images and generating diverse views over time is challenging due to scene complexity and temporal dynamics. Despite advancements in neural implicit models, limitations persist: (i) Inadequate Scene…

Computer Vision and Pattern Recognition · Computer Science 2024-02-23 Zeyu Yang , Hongye Yang , Zijie Pan , Li Zhang

We address the problem of recovering a time-varying 4D distribution from a sparse sequence of 2D projections - analogous to novel-view synthesis from sparse cameras, but applied to the 4D transverse phase space density $\rho(x,p_x,y,p_y)$…

Accelerator Physics · Physics 2026-04-08 Alexander Scheinker , Alexander Plastun , Peter Ostroumov

Image generation models trained on large datasets can synthesize high-quality images but often produce spatially inconsistent and distorted images due to limited information about the underlying structures and spatial layouts. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Hyundo Lee , Suhyung Choi , Inwoo Hwang , Byoung-Tak Zhang

Recent 3D large reconstruction models typically employ a two-stage process, including first generate multi-view images by a multi-view diffusion model, and then utilize a feed-forward model to reconstruct images to 3D content.However,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Zhenyu Tang , Junwu Zhang , Xinhua Cheng , Wangbo Yu , Chaoran Feng , Yatian Pang , Bin Lin , Li Yuan

Designing effective camera trajectories in virtual 3D environments is a challenging task even for experienced animators. Despite an elaborate film grammar, forged through years of experience, that enables the specification of camera motions…

Graphics · Computer Science 2024-02-27 Hongda Jiang , Xi Wang , Marc Christie , Libin Liu , Baoquan Chen

Diffusion models have emerged as effective tools for generating diverse and high-quality content. However, their capability in high-resolution image generation, particularly for panoramic images, still faces challenges such as visible seams…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Teng Zhou , Yongchuan Tang