中文
相关论文

相关论文: DynamicScaler: Seamless and Scalable Video Generat…

200 篇论文

Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are generated earlier in the denoising process while high-frequency details emerge only in later…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Howard Xiao , Brian Chao , Lior Yariv , Gordon Wetzstein

State-of-the-art video generation models produce remarkable photorealism, but they lack the precise control required to align generated content with specific scene requirements. Furthermore, without an underlying explicit geometry, these…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dana Cohen-Bar , Ido Sobol , Raphael Bensadoun , Shelly Sheynin , Oran Gafni , Or Patashnik , Daniel Cohen-Or , Amit Zohar

Image tokenization plays a central role in modern generative modeling by mapping visual inputs into compact representations that serve as an intermediate signal between pixels and generative models. Diffusion-based decoders have recently…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Chuhan Wang , Hao Chen

Generative world models have become essential data engines for autonomous driving, yet most existing efforts focus on videos or occupancy grids, overlooking the unique LiDAR properties. Extending LiDAR generation to dynamic 4D world…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Ao Liang , Youquan Liu , Yu Yang , Dongyue Lu , Linfeng Li , Lingdong Kong , Huaici Zhao , Wei Tsang Ooi

In this paper, we introduce \textbf{DimensionX}, a framework designed to generate photorealistic 3D and 4D scenes from just a single image with video diffusion. Our approach begins with the insight that both the spatial structure of a 3D…

计算机视觉与模式识别 · 计算机科学 2024-11-08 Wenqiang Sun , Shuo Chen , Fangfu Liu , Zilong Chen , Yueqi Duan , Jun Zhang , Yikai Wang

Scene-level 3D generation represents a critical frontier in multimedia and computer graphics, yet existing approaches either suffer from limited object categories or lack editing flexibility for interactive applications. In this paper, we…

图形学 · 计算机科学 2025-04-18 Wenqi Dong , Bangbang Yang , Zesong Yang , Yuan Li , Tao Hu , Hujun Bao , Yuewen Ma , Zhaopeng Cui

Text-driven 3D scene generation has seen significant advancements recently. However, most existing methods generate single-view images using generative models and then stitch them together in 3D space. This independent generation for each…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Wenrui Li , Fucheng Cai , Yapeng Mi , Zhe Yang , Wangmeng Zuo , Xingtao Wang , Xiaopeng Fan

With the advent of portable 360{\deg} cameras, panorama has gained significant attention in applications like virtual reality (VR), virtual tours, robotics, and autonomous driving. As a result, wide-baseline panorama view synthesis has…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Cheng Zhang , Haofei Xu , Qianyi Wu , Camilo Cruz Gambardella , Dinh Phung , Jianfei Cai

Despite remarkable achievements in video synthesis, achieving granular control over complex dynamics, such as nuanced movement among multiple interacting objects, still presents a significant hurdle for dynamic world modeling, compounded by…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Pengxiang Li , Kai Chen , Zhili Liu , Ruiyuan Gao , Lanqing Hong , Guo Zhou , Hua Yao , Dit-Yan Yeung , Huchuan Lu , Xu Jia

Realistic reconstruction of dynamic 4D scenes from monocular videos is essential for understanding the physical world. Despite recent progress in neural rendering, existing methods often struggle to recover accurate 3D geometry and…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Haoran Zhou , Gim Hee Lee

Recent advances in camera-controllable video generation have been constrained by the reliance on static-scene datasets with relative-scale camera annotations, such as RealEstate10K. While these datasets enable basic viewpoint control, they…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Guangcong Zheng , Teng Li , Xianpan Zhou , Xi Li

Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack of high-resolution data and constrained computation resources, hampering their ability to generate high-fidelity…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Haonan Qiu , Ning Yu , Ziqi Huang , Paul Debevec , Ziwei Liu

We propose a decoupled 3D scene generation framework called SceneMaker in this work. Due to the lack of sufficient open-set de-occlusion and pose estimation priors, existing methods struggle to simultaneously produce high-quality geometry…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Yukai Shi , Weiyu Li , Zihao Wang , Hongyang Li , Xingyu Chen , Ping Tan , Lei Zhang

We address the problem of synthesizing novel views from a monocular video depicting a complex dynamic scene. State-of-the-art methods based on temporally varying Neural Radiance Fields (aka dynamic NeRFs) have shown impressive results on…

计算机视觉与模式识别 · 计算机科学 2023-04-25 Zhengqi Li , Qianqian Wang , Forrester Cole , Richard Tucker , Noah Snavely

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Ziqi Cai , Taoyu Yang , Zheng Chang , Si Li , Han Jiang , Shuchen Weng , Boxin Shi

Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus on character consistency, they often overlook the crucial…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Quanjian Song , Donghao Zhou , Jingyu Lin , Fei Shen , Jiaze Wang , Xiaowei Hu , Cunjian Chen , Pheng-Ann Heng

Scalability in terms of object density in a scene is a primary challenge in unsupervised sequential object-oriented representation learning. Most of the previous models have been shown to work only on scenes with a few objects. In this…

机器学习 · 计算机科学 2020-03-06 Jindong Jiang , Sepehr Janghorbani , Gerard de Melo , Sungjin Ahn

Generative models have gained significant attention in novel view synthesis (NVS) by alleviating the reliance on dense multi-view captures. However, existing methods typically fall into a conventional paradigm, where generative models first…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Weiliang Chen , Jiayi Bi , Yuanhui Huang , Wenzhao Zheng , Yueqi Duan

We introduce Diff4Splat, a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Panwang Pan , Chenguo Lin , Jingjing Zhao , Chenxin Li , Yuchen Lin , Haopeng Li , Honglei Yan , Kairun Wen , Yunlong Lin , Yixuan Yuan , Yadong Mu

This paper aims to tackle the problem of photorealistic view synthesis from vehicle sensor data. Recent advancements in neural scene representation have achieved notable success in rendering high-quality autonomous driving scenes, but the…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Yunzhi Yan , Zhen Xu , Haotong Lin , Haian Jin , Haoyu Guo , Yida Wang , Kun Zhan , Xianpeng Lang , Hujun Bao , Xiaowei Zhou , Sida Peng