English
Related papers

Related papers: Layout-your-3D: Controllable and Precise 3D Genera…

200 papers

Recent advancements in 3D scene editing have been propelled by the rapid development of generative models. Existing methods typically utilize generative models to perform text-guided editing on 3D representations, such as 3D Gaussian…

Graphics · Computer Science 2025-05-27 Yansong Qu , Dian Chen , Xinyang Li , Xiaofan Li , Shengchuan Zhang , Liujuan Cao , Rongrong Ji

We present DIRECT-3D, a diffusion-based 3D generative model for creating high-quality 3D assets (represented by Neural Radiance Fields) from text prompts. Unlike recent 3D generative models that rely on clean and well-aligned 3D data,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Qihao Liu , Yi Zhang , Song Bai , Adam Kortylewski , Alan Yuille

Manually creating 3D environments for AR/VR applications is a complex process requiring expert knowledge in 3D modeling software. Pioneering works facilitate this process by generating room meshes conditioned on textual style descriptions.…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Jonas Schult , Sam Tsai , Lukas Höllein , Bichen Wu , Jialiang Wang , Chih-Yao Ma , Kunpeng Li , Xiaofang Wang , Felix Wimbauer , Zijian He , Peizhao Zhang , Bastian Leibe , Peter Vajda , Ji Hou

If a picture is worth thousand words, a moving 3d shape must be worth a million. We build upon the success of recent generative methods that create images fitting the semantics of a text prompt, and extend it to the controlled generation of…

Machine Learning · Computer Science 2021-09-28 Nikolay Jetchev

We present GuidedSceneGen, a text-to-3D generation framework that produces metrically accurate, globally consistent, and semantically interpretable indoor scenes. Unlike prior text-driven methods that often suffer from geometric drift or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Stefan Ainetter , Thomas Deixelberger , Edoardo A. Dominici , Philipp Drescher , Konstantinos Vardis , Markus Steinberger

Recent works on text-to-3d generation show that using only 2D diffusion supervision for 3D generation tends to produce results with inconsistent appearances (e.g., faces on the back view) and inaccurate shapes (e.g., animals with extra…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Cheng Chen , Xiaofeng Yang , Fan Yang , Chengzeng Feng , Zhoujie Fu , Chuan-Sheng Foo , Guosheng Lin , Fayao Liu

Recent approaches have achieved great success in image generation from structured inputs, e.g., semantic segmentation, scene graph or layout. Although these methods allow specification of objects and their locations at image-level, they…

Computer Vision and Pattern Recognition · Computer Science 2020-08-28 Ke Ma , Bo Zhao , Leonid Sigal

Generating 3D shapes at part level is pivotal for downstream applications such as mesh retopology, UV mapping, and 3D printing. However, existing part-based generation methods often lack sufficient controllability and suffer from poor…

Generating 3D scenes from natural language holds great promise for applications in gaming, film, and design. However, existing methods struggle with automation, 3D consistency, and fine-grained control. We present DreamScene, an end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Haoran Li , Yuli Tian , Kun Lan , Yong Liao , Lin Wang , Pan Hui , Peng Yuan Zhou

Recent advancements in text-to-3D generation have shown remarkable results by leveraging 3D priors in combination with 2D diffusion. However, previous methods utilize 3D priors that lack detailed and complex structural information, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-11-05 Hao Wu , Hao Wang , Ruochong Li , Xuran Ma , Hui Xiong

We present Turbo3D, an ultra-fast text-to-3D system capable of generating high-quality Gaussian splatting assets in under one second. Turbo3D employs a rapid 4-step, 4-view diffusion generator and an efficient feed-forward Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Hanzhe Hu , Tianwei Yin , Fujun Luan , Yiwei Hu , Hao Tan , Zexiang Xu , Sai Bi , Shubham Tulsiani , Kai Zhang

Recent multimodal large language models have achieved strong performance in unified text and image understanding and generation, yet extending such native capability to 3D remains challenging due to limited data. Compared to abundant 2D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Chongjie Ye , Cheng Cao , Chuanyu Pan , Yiming Hao , Yihao Zhi , Yuanming Hu , Xiaoguang Han

Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for 3D scene editing still exhibit certain shortcomings,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Shuangkang Fang , Yufeng Wang , Yi-Hsuan Tsai , Yi Yang , Wenrui Ding , Shuchang Zhou , Ming-Hsuan Yang

Current controls over diffusion models (e.g., through text or ControlNet) for image generation fall short in recognizing abstract, continuous attributes like illumination direction or non-rigid shape change. In this paper, we present an…

Computer Vision and Pattern Recognition · Computer Science 2024-02-14 Ta-Ying Cheng , Matheus Gadelha , Thibault Groueix , Matthew Fisher , Radomir Mech , Andrew Markham , Niki Trigoni

Realistic 3D human generation from text prompts is a desirable yet challenging task. Existing methods optimize 3D representations like mesh or neural fields via score distillation sampling (SDS), which suffers from inadequate fine details…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Xian Liu , Xiaohang Zhan , Jiaxiang Tang , Ying Shan , Gang Zeng , Dahua Lin , Xihui Liu , Ziwei Liu

Recently, the text-to-3D task has developed rapidly due to the appearance of the SDS method. However, the SDS method always generates 3D objects with poor quality due to the over-smooth issue. This issue is attributed to two factors: 1) the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Yiming Zhong , Xiaolin Zhang , Yao Zhao , Yunchao Wei

Automatic furniture layout is long desired for convenient interior design. Leveraging the remarkable visual reasoning capabilities of multimodal large language models (MLLMs), recent methods address layout generation in a static manner,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-01 Can Wang , Hongliang Zhong , Menglei Chai , Mingming He , Dongdong Chen , Jing Liao

We introduce PAT3D, the first physics-augmented text-to-3D scene generation framework that integrates vision-language models with physics-based simulation to produce physically plausible, simulation-ready, and intersection-free 3D scenes.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Guying Lin , Kemeng Huang , Michael Liu , Ruihan Gao , Hanke Chen , Lyuhao Chen , Beijia Lu , Taku Komura , Yuan Liu , Jun-Yan Zhu , Minchen Li

Graphic design forms the cornerstone of modern visual communication, serving as a vital medium for promoting cultural and commercial events. Recent advances have explored automating this process using Large Multimodal Models (LMMs), yet…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Jiazhe Wei , Ken Li , Tianyu Lao , Haofan Wang , Liang Wang , Caifeng Shan , Chenyang Si

Although recent 3D-native generators have made great progress in synthesizing reliable geometry, they still fall short in achieving realistic appearances. A key obstacle lies in the lack of diverse and high-quality real-world 3D assets with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xinyue Liang , Zhinyuan Ma , Lingchen Sun , Yanjun Guo , Lei Zhang