中文
相关论文

相关论文: GALA3D: Towards Text-to-3D Complex Scene Generatio…

200 篇论文

In recent years, 3D generation has made great strides in both academia and industry. However, generating 3D scenes from a single RGB image remains a significant challenge, as current approaches often struggle to ensure both object…

图形学 · 计算机科学 2026-02-18 Xiang Tang , Ruotong Li , Xiaopeng Fan

Recent advancements in Generalizable Gaussian Splatting have enabled robust 3D reconstruction from sparse input views by utilizing feed-forward Gaussian Splatting models, achieving superior cross-scene generalization. However, while many…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Zhicong Wu , Hongbin Xu , Gang Xu , Ping Nie , Zhixin Yan , Jinkai Zheng , Liangqiong Qu , Ming Li , Liqiang Nie

We introduce Text2Immersion, an elegant method for producing high-quality 3D immersive scenes from text prompts. Our proposed pipeline initiates by progressively generating a Gaussian cloud using pre-trained 2D diffusion and depth…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Hao Ouyang , Kathryn Heal , Stephen Lombardi , Tiancheng Sun

Generating realistic 3D scenes from text is crucial for immersive applications like VR, AR, and gaming. While text-driven approaches promise efficiency, existing methods suffer from limited 3D-text data and inconsistent multi-view…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Xin Zhang , Shen Chen , Jiale Zhou , Lei Li

Generating and editing a 3D scene guided by natural language poses a challenge, primarily due to the complexity of specifying the positional relations and volumetric changes within the 3D space. Recent advancements in Large Language Models…

计算机视觉与模式识别 · 计算机科学 2023-05-26 Yiqi Lin , Hao Wu , Ruichen Wang , Haonan Lu , Xiaodong Lin , Hui Xiong , Lin Wang

With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Xiaolu Hou , Mingcheng Li , Dingkang Yang , Jiawei Chen , Ziyun Qian , Xiao Zhao , Yue Jiang , Jinjie Wei , Qingyao Xu , Lihua Zhang

Text-to-3D scene generation holds immense potential for the gaming, film, and architecture sectors. Despite significant progress, existing methods struggle with maintaining high quality, consistency, and editing flexibility. In this paper,…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Haoran Li , Haolin Shi , Wenli Zhang , Wenjun Wu , Yong Liao , Lin Wang , Lik-hang Lee , Pengyuan Zhou

The creation of high-quality 3D assets is paramount for applications in digital heritage preservation, entertainment, and robotics. Traditionally, this process necessitates skilled professionals and specialized software for the modeling,…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Shen Chen , Jiale Zhou , Zhongyu Jiang , Tianfang Zhang , Zongkai Wu , Jenq-Neng Hwang , Lei Li

Recent advances in 3D content creation mostly leverage optimization-based 3D generation via score distillation sampling (SDS). Though promising results have been exhibited, these methods often suffer from slow per-sample optimization,…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Jiaxiang Tang , Jiawei Ren , Hang Zhou , Ziwei Liu , Gang Zeng

In this work, we introduce Prometheus, a 3D-aware latent diffusion model for text-to-3D generation at both object and scene levels in seconds. We formulate 3D scene generation as multi-view, feed-forward, pixel-aligned 3D Gaussian…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Yuanbo Yang , Jiahao Shao , Xinyang Li , Yujun Shen , Andreas Geiger , Yiyi Liao

Most recent advances in 3D generative modeling rely on diffusion or flow-matching formulations. We instead explore a fully autoregressive alternative and introduce GaussianGPT, a transformer-based model that directly generates 3D Gaussians…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Nicolas von Lützow , Barbara Rössle , Katharina Schmid , Matthias Nießner

We present GuidedSceneGen, a text-to-3D generation framework that produces metrically accurate, globally consistent, and semantically interpretable indoor scenes. Unlike prior text-driven methods that often suffer from geometric drift or…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Stefan Ainetter , Thomas Deixelberger , Edoardo A. Dominici , Philipp Drescher , Konstantinos Vardis , Markus Steinberger

3D Gaussian Splatting (3DGS) enables the reconstruction of intricate digital 3D assets from multi-view images by leveraging a set of 3D Gaussian primitives for rendering. Its explicit and discrete representation facilitates the seamless…

图形学 · 计算机科学 2025-05-13 Xijie Yang , Linning Xu , Lihan Jiang , Dahua Lin , Bo Dai

We propose a method to enhance 3D Gaussian Splatting (3DGS)~\cite{Kerbl2023}, addressing challenges in initialization, optimization, and density control. Gaussian Splatting is an alternative for rendering realistic images while supporting…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Xingjun Wang , Lianlei Shan

We introduce GaussianZoom, a generative zoom-in 3D reconstruction system with an iterative progressive framework that combines geometry-consistent scene modeling and multi-scale semantic reasoning to enable high-fidelity extreme zoom-in…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Jiale Shi , Jiarui Hu , Zesong Yang , Kaixuan Luan , Hujun Bao , Zhaopeng Cui

State-of-the-art novel view synthesis methods achieve impressive results for multi-view captures of static 3D scenes. However, the reconstructed scenes still lack "liveliness," a key component for creating engaging 3D experiences. Recently,…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Thomas Wimmer , Michael Oechsle , Michael Niemeyer , Federico Tombari

Despite recent advancements in neural 3D reconstruction, the dependence on dense multi-view captures restricts their broader applicability. Additionally, 3D scene generation is vital for advancing embodied AI and world models, which depend…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Yuxin Zhang , Ziyu Lu , Hongbo Duan , Keyu Fan , Pengting Luo , Peiyu Zhuang , Mengyu Yang , Houde Liu

We present Layout-Your-3D, a framework that allows controllable and compositional 3D generation from text prompts. Existing text-to-3D methods often struggle to generate assets with plausible object interactions or require tedious…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Junwei Zhou , Xueting Li , Lu Qi , Ming-Hsuan Yang

Designing complex 3D scenes has been a tedious, manual process requiring domain expertise. Emerging text-to-3D generative models show great promise for making this task more intuitive, but existing approaches are limited to object-level…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Ryan Po , Gordon Wetzstein

Empowering 3D Gaussian Splatting with generalization ability is appealing. However, existing generalizable 3D Gaussian Splatting methods are largely confined to narrow-range interpolation between stereo images due to their heavy backbones,…

计算机视觉与模式识别 · 计算机科学 2024-10-30 Yunsong Wang , Tianxin Huang , Hanlin Chen , Gim Hee Lee