English
Related papers

Related papers: Spatia: Video Generation with Updatable Spatial Me…

200 papers

Reconstructing dynamic 3D scenes from monocular video remains fundamentally challenging due to the need to jointly infer motion, structure, and appearance from limited observations. Existing dynamic scene reconstruction methods based on…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Jiahui Li , Shengeng Tang , Jingxuan He , Gang Huang , Zhangye Wang , Yantao Pan , Lechao Cheng

The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation and feature…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jiamin Wang , Yichen Yao , Xiang Feng , Hang Wu , Yaming Wang , Qingqiu Huang , Yuexin Ma , Xinge Zhu

Rapid advances in the field of generative AI and text-to-image methods in particular have transformed the way we interact with and perceive computer-generated imagery today. In parallel, much progress has been made in 3D face…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Mirela Ostrek , Justus Thies

While open-source video generation and editing models have made significant progress, individual models are typically limited to specific tasks, failing to meet the diverse needs of users. Effectively coordinating these models can unlock a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Rong-Cheng Tu , Wenhao Sun , Zhao Jin , Jingyi Liao , Jiaxing Huang , Dacheng Tao

Video spatial reasoning requires accumulating viewpoint-dependent evidence over time while retaining information useful to the question being asked. Existing spatial video-language models improve geometric perception and long-range context…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Xianqiang Gao , Qizhi Chen , Delin Qu , Haoming Song , Zhigang Wang , Bin Zhao , Dong Wang , Xuelong Li

Recent advancements in 3D generation are predominantly propelled by improvements in 3D-aware image diffusion models. These models are pretrained on Internet-scale image data and fine-tuned on massive 3D data, offering the capability of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Zeyu Yang , Zijie Pan , Chun Gu , Li Zhang

This paper explores the innovative application of Stable Video Diffusion (SVD), a diffusion model that revolutionizes the creation of dynamic video content from static images. As digital media and design industries accelerate, SVD emerges…

Human-Computer Interaction · Computer Science 2024-05-24 Elijah Miller , Thomas Dupont , Mingming Wang

Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Shangchen Zhou , Peiqing Yang , Jianyi Wang , Yihang Luo , Chen Change Loy

Understanding how visual information is encoded in biological and artificial systems often requires vision scientists to generate appropriate stimuli to test specific hypotheses. Although deep neural network models have revolutionized the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Antonino Greco , Markus Siegel

Autoregressive video diffusion models have proved effective for world modeling and interactive scene generation, with Minecraft gameplay as a representative application. To faithfully simulate play, a model must generate natural content…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Junchao Huang , Xinting Hu , Boyao Han , Shaoshuai Shi , Zhuotao Tian , Tianyu He , Li Jiang

Realistic reconstruction of dynamic 4D scenes from monocular videos is essential for understanding the physical world. Despite recent progress in neural rendering, existing methods often struggle to recover accurate 3D geometry and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Haoran Zhou , Gim Hee Lee

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Chin-Yang Lin , Cheng Sun , Fu-En Yang , Min-Hung Chen , Yen-Yu Lin , Yu-Lun Liu

In the field of action recognition, video clips are always treated as ordered frames for subsequent processing. To achieve spatio-temporal perception, existing approaches propose to embed adjacent temporal interaction in the convolutional…

Computer Vision and Pattern Recognition · Computer Science 2022-02-01 Rongchang Li , Xiao-Jun Wu , Tianyang Xu

Novel view synthesis (NVS) boosts immersive experiences in computer vision and graphics. Existing techniques, though progressed, rely on dense multi-view observations, restricting their application. This work takes on the challenge of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Songchun Zhang , Huiyao Xu , Sitong Guo , Zhongwei Xie , Hujun Bao , Weiwei Xu , Changqing Zou

The ultimate goal of video generation is to satisfy a fundamental trilemma: achieving high visual quality, maintaining rigorous physical consistency, and enabling precise controllability. While recent models can maintain this balance in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Tianshuo Xu , Zhifei Chen , Leyi Wu , Hao Lu , Ying-cong Chen

While 3D content generation has advanced significantly, existing methods still face challenges with input formats, latent space design, and output representations. This paper introduces a novel 3D generation framework that addresses these…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Yushi Lan , Shangchen Zhou , Zhaoyang Lyu , Fangzhou Hong , Shuai Yang , Bo Dai , Xingang Pan , Chen Change Loy

Creating deformable 3D content has gained increasing attention with the rise of text-to-image and image-to-video generative models. While these models provide rich semantic priors for appearance, they struggle to capture the physical…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Jixuan He , Chieh Hubert Lin , Lu Qi , Ming-Hsuan Yang

For recent diffusion-based generative models, maintaining consistent content across a series of generated images, especially those containing subjects and complex details, presents a significant challenge. In this paper, we propose a new…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yupeng Zhou , Daquan Zhou , Ming-Ming Cheng , Jiashi Feng , Qibin Hou

Novel View Synthesis plays a crucial role by generating new 2D renderings from multi-view images of 3D scenes. However, capturing high-speed scenes with conventional cameras often leads to motion blur, hindering the effectiveness of 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Jiyuan Zhang , Kang Chen , Shiyan Chen , Yajing Zheng , Tiejun Huang , Zhaofei Yu

We propose Camera Splatting, a novel view optimization framework for novel view synthesis. Each camera is modeled as a 3D Gaussian, referred to as a camera splat, and virtual cameras, termed point cameras, are placed at 3D points sampled…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Gahye Lee , Hyomin Kim , Gwangjin Ju , Jooeun Son , Hyejeong Yoon , Seungyong Lee