中文
相关论文

相关论文: GEN3C: 3D-Informed World-Consistent Video Generati…

200 篇论文

Recent advances in foundational Video Diffusion Models (VDMs) have yielded significant progress. Yet, despite the remarkable visual quality of generated videos, reconstructing consistent 3D scenes from these outputs remains challenging, due…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Yisu Zhang , Chenjie Cao , Tengfei Wang , Xuhui Zuo , Junta Wu , Jianke Zhu , Chunchao Guo

Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditions to guide generation. However, these models primarily animate static images without…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Luis Denninger , Sina Mokhtarzadeh Azar , Juergen Gall

Existing point cloud completion methods, which typically depend on predefined synthetic training datasets, encounter significant challenges when applied to out-of-distribution, real-world scans. To overcome this limitation, we introduce a…

计算机视觉与模式识别 · 计算机科学 2025-02-28 An Li , Zhe Zhu , Mingqiang Wei

Recent generative models can produce high-fidelity videos, yet they often exhibit 3D spatial geometric inconsistencies. Existing evaluation methods fail to accurately characterize these inconsistencies: fidelity-centric metrics like FVD are…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Weijia Dou , Wenzhao Zheng , Weiliang Chen , Yu Zheng , Jie Zhou , Jiwen Lu

We extend multimodal transformers to include 3D camera motion as a conditioning signal for the task of video generation. Generative video models are becoming increasingly powerful, thus focusing research efforts on methods of controlling…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Andrew Marmon , Grant Schindler , José Lezama , Dan Kondratyuk , Bryan Seybold , Irfan Essa

Recent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges,…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Teng Li , Guangcong Zheng , Rui Jiang , Shuigen Zhan , Tao Wu , Yehao Lu , Yining Lin , Chuanyun Deng , Yepan Xiong , Min Chen , Lin Cheng , Xi Li

Recently, breakthroughs in video modeling have allowed for controllable camera trajectories in generated videos. However, these methods cannot be directly applied to user-provided videos that are not generated by a video model. In this…

计算机视觉与模式识别 · 计算机科学 2024-11-08 David Junhao Zhang , Roni Paiss , Shiran Zada , Nikhil Karnad , David E. Jacobs , Yael Pritch , Inbar Mosseri , Mike Zheng Shou , Neal Wadhwa , Nataniel Ruiz

A recent frontier in computer vision has been the task of 3D video generation, which consists of generating a time-varying 3D representation of a scene. To generate dynamic 3D scenes, current methods explicitly model 3D temporal dynamics by…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Rishab Parthasarathy , Zachary Ankner , Aaron Gokaslan

World models based on video generation demonstrate remarkable potential for simulating interactive environments but face persistent difficulties in two key areas: maintaining long-term content consistency when scenes are revisited and…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Tianxing Xu , Zixuan Wang , Guangyuan Wang , Li Hu , Zhongyi Zhang , Peng Zhang , Bang Zhang , Song-Hai Zhang

Despite remarkable progress in video generation, maintaining long-term scene consistency upon revisiting previously explored areas remains challenging. Existing solutions rely either on explicitly constructing 3D geometry, which suffers…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Jia Li , Han Yan , Yihang Chen , Siqi Li , Xibin Song , Yifu Wang , Jianfei Cai , Tien-Tsin Wong , Pan Ji

We introduce VIVE3D, a novel approach that extends the capabilities of image-based 3D GANs to video editing and is able to represent the input video in an identity-preserving and temporally consistent way. We propose two new building…

计算机视觉与模式识别 · 计算机科学 2023-03-29 Anna Frühstück , Nikolaos Sarafianos , Yuanlu Xu , Peter Wonka , Tony Tung

Controllable generative models for images and videos have seen significant success, yet 3D scene generation, especially in unbounded scenarios like autonomous driving, remains underdeveloped. Existing methods lack flexible controllability…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Ruiyuan Gao , Kai Chen , Zhihao Li , Lanqing Hong , Zhenguo Li , Qiang Xu

Accurate reconstruction of complex dynamic scenes from just a single viewpoint continues to be a challenging task in computer vision. Current dynamic novel view synthesis methods typically require videos from many different camera…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Basile Van Hoorick , Rundi Wu , Ege Ozguroglu , Kyle Sargent , Ruoshi Liu , Pavel Tokmakov , Achal Dave , Changxi Zheng , Carl Vondrick

Tremendous progress in deep generative models has led to photorealistic image synthesis. While achieving compelling results, most approaches operate in the two-dimensional image domain, ignoring the three-dimensional nature of our world.…

计算机视觉与模式识别 · 计算机科学 2021-04-01 Michael Niemeyer , Andreas Geiger

We present 3DScenePrompt, a framework that generates the next video chunk from arbitrary-length input while enabling precise camera control and preserving scene consistency. Unlike methods conditioned on a single image or a short clip, we…

计算机视觉与模式识别 · 计算机科学 2025-12-16 JoungBin Lee , Jaewoo Jung , Jisang Han , Takuya Narihira , Kazumi Fukuda , Junyoung Seo , Sunghwan Hong , Yuki Mitsufuji , Seungryong Kim

The creation of manufacturable and editable 3D shapes through Computer-Aided Design (CAD) remains a highly manual and time-consuming task, hampered by the complex topology of boundary representations of 3D solids and unintuitive design…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Md Ferdous Alam , Faez Ahmed

Single-image 3D generation has emerged as a prominent research topic, playing a vital role in virtual reality, 3D modeling, and digital content creation. However, existing methods face challenges such as a lack of multi-view geometric…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Jinbo Yan , Alan Zhao , Yixin Hu

We present IDC-Net (Image-Depth Consistency Network), a novel framework designed to generate RGB-D video sequences under explicit camera trajectory control. Unlike approaches that treat RGB and depth generation separately, IDC-Net jointly…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Lijuan Liu , Wenfa Li , Dongbo Zhang , Shuo Wang , Shaohui Jiao

The creation of 3D human face avatars from a single unconstrained image is a fundamental task that underlies numerous real-world vision and graphics applications. Despite the significant progress made in generative models, existing methods…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Wenqing Wang , Haosen Yang , Josef Kittler , Xiatian Zhu

Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise camera control with global scene understanding in text-to-image generation by learning…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Xinxuan Lu , Charless Fowlkes , Alexander C. Berg