English
Related papers

Related papers: VideoFrom3D: 3D Scene Video Generation via Complem…

200 papers

Recent advancements in 3D generation are predominantly propelled by improvements in 3D-aware image diffusion models. These models are pretrained on Internet-scale image data and fine-tuned on massive 3D data, offering the capability of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Zeyu Yang , Zijie Pan , Chun Gu , Li Zhang

Generating multi-view images based on text or single-image prompts is a critical capability for the creation of 3D content. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Qi Zuo , Xiaodong Gu , Lingteng Qiu , Yuan Dong , Zhengyi Zhao , Weihao Yuan , Rui Peng , Siyu Zhu , Zilong Dong , Liefeng Bo , Qixing Huang

The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. However, existing approaches struggle to infer massive amounts of missing geometry in large…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Dehui Wang , Congsheng Xu , Rong Wei , Yue Shi , Shoufa Chen , Dingxiang Luo , Tianshuo Yang , Xiaokang Yang , Wei Sui , Yusen Qin , Rui Tang , Yao Mu

How can one efficiently generate high-quality, wide-scope 3D scenes from arbitrary single images? Existing methods suffer several drawbacks, such as requiring multi-view data, time-consuming per-scene optimization, distorted geometry in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Hanwen Liang , Junli Cao , Vidit Goel , Guocheng Qian , Sergei Korolev , Demetri Terzopoulos , Konstantinos N. Plataniotis , Sergey Tulyakov , Jian Ren

3D AI-generated content (AIGC) has made it increasingly accessible for anyone to become a 3D content creator. While recent methods leverage Score Distillation Sampling to distill 3D objects from pretrained image diffusion models, they often…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Yuxiao Yang , Peihao Li , Yuhong Zhang , Junzhe Lu , Xianglong He , Minghan Qin , Weitao Wang , Haoqian Wang

This paper presents a novel method for building scalable 3D generative models utilizing pre-trained video diffusion models. The primary obstacle in developing foundation 3D generative models is the limited availability of 3D data. Unlike…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Junlin Han , Filippos Kokkinos , Philip Torr

The ability to generate virtual environments is crucial for applications ranging from gaming to physical AI domains such as robotics, autonomous driving, and industrial AI. Current learning-based 3D reconstruction methods rely on the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Sherwin Bahmani , Tianchang Shen , Jiawei Ren , Jiahui Huang , Yifeng Jiang , Haithem Turki , Andrea Tagliasacchi , David B. Lindell , Zan Gojcic , Sanja Fidler , Huan Ling , Jun Gao , Xuanchi Ren

Creating flexible 3D scenes from a single image is vital when direct 3D data acquisition is costly or impractical. We introduce NavCrafter, a novel framework that explores 3D scenes from a single image by synthesizing novel-view video…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Hongbo Duan , Peiyu Zhuang , Yi Liu , Zhengyang Zhang , Yuxin Zhang , Pengting Luo , Fangming Liu , Xueqian Wang

Previous works leveraging video models for image-to-3D scene generation tend to suffer from geometric distortions and blurry content. In this paper, we renovate the pipeline of image-to-3D scene generation by unlocking the potential of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yuhao Wan , Lijuan Liu , Jingzhi Zhou , Zihan Zhou , Xuying Zhang , Dongbo Zhang , Shaohui Jiao , Qibin Hou , Ming-Ming Cheng

Mesh models have become increasingly accessible for numerous cities; however, the lack of realistic textures restricts their application in virtual urban navigation and autonomous driving. To address this, this paper proposes MeSS…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Xuyang Chen , Zhijun Zhai , Kaixuan Zhou , Zengmao Wang , Jianan He , Dong Wang , Yanfeng Zhang , mingwei Sun , Rüdiger Westermann , Konrad Schindler , Liqiu Meng

Despite recent successes in novel view synthesis using 3D Gaussian Splatting (3DGS), modeling scenes with sparse inputs remains a challenge. In this work, we address two critical yet overlooked issues in real-world sparse-input modeling:…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Yingji Zhong , Zhihao Li , Dave Zhenyu Chen , Lanqing Hong , Dan Xu

Generating 4D scenes from a single-view video is inherently ill-posed: a single viewpoint lacks the information needed to recover a complete, dynamic scene with full coverage. Existing methods are typically limited to monocular videos,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Tingxi Chen , Ke Hao , Yabo Chen , Zhengxue Cheng , Rong Xie , Li Song , Haibin Huang , Chi Zhang , Xuelong Li

Video generation models have demonstrated great capabilities of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free approach for generating…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Peng Dai , Feitong Tan , Qiangeng Xu , David Futschik , Ruofei Du , Sean Fanello , Xiaojuan Qi , Yinda Zhang

Sparse-view 3D reconstruction is essential for modeling scenes from casual captures, but remain challenging for non-generative reconstruction. Existing diffusion-based approaches mitigates this issues by synthesizing novel views, but they…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Yutian Chen , Shi Guo , Renbiao Jin , Tianshuo Yang , Xin Cai , Yawen Luo , Mingxin Yang , Mulin Yu , Linning Xu , Tianfan Xue

Geometric foundation models show promise in 3D reconstruction, yet their progress is severely constrained by the scarcity of diverse, large-scale 3D annotations. While Internet videos offer virtually unlimited raw data, utilizing them as a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Zihui Gao , Ke Liu , Donny Y. Chen , Duochao Shi , Guosheng Lin , Hao Chen , Chunhua Shen

Generating complete digital twins from videos requires precise camera control, global scene coverage, and strict spatial-temporal consistency constraints that remain challenging for perspective video generators due to their limited field of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Ting-Hsuan Chen , Ying-Huan Chen , Tao Tu , Jie-Ying Lee , Cho-Ying Wu , Fangzhou Lin , Hengyuan Zhang , David Paz , Xinyu Huang , Yuliang Guo , Yu-Lun Liu , Yue Wang , Liu Ren

Panoramic image stitching provides a unified, wide-angle view of a scene that extends beyond the camera's field of view. Stitching frames of a panning video into a panoramic photograph is a well-understood problem for stationary scenes, but…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Jingwei Ma , Erika Lu , Roni Paiss , Shiran Zada , Aleksander Holynski , Tali Dekel , Brian Curless , Michael Rubinstein , Forrester Cole

In this paper, we tackle a new task of 3D object synthesis, where a 3D model is composited with another object category to create a novel 3D model. However, most existing text/image/3D-to-3D methods struggle to effectively integrate…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Zeren Xiong , Zikun Chen , Zedong Zhang , Xiang Li , Ying Tai , Jian Yang , Jun Li

Image-based 3D generation has vast applications in robotics and gaming, where high-quality, diverse outputs and consistent 3D representations are crucial. However, existing methods have limitations: 3D diffusion models are limited by…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Ye Tao , Jiawei Zhang , Yahao Shi , Dongqing Zou , Bin Zhou

We present GuidedSceneGen, a text-to-3D generation framework that produces metrically accurate, globally consistent, and semantically interpretable indoor scenes. Unlike prior text-driven methods that often suffer from geometric drift or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Stefan Ainetter , Thomas Deixelberger , Edoardo A. Dominici , Philipp Drescher , Konstantinos Vardis , Markus Steinberger