English
Related papers

Related papers: Text2Immersion: Generative Immersive Scene with 3D…

200 papers

As multimodal language models advance, their application to 3D scene understanding is a fast-growing frontier, driving the development of 3D Vision-Language Models (VLMs). Current methods show strong dependence on object detectors,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Anna-Maria Halacheva , Jan-Nico Zaech , Xi Wang , Danda Pani Paudel , Luc Van Gool

Recently, Gaussian Splatting, a method that represents a 3D scene as a collection of Gaussian distributions, has gained significant attention in addressing the task of novel view synthesis. In this paper, we highlight a fundamental…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Haoxuan Qu , Zhuoling Li , Hossein Rahmani , Yujun Cai , Jun Liu

We introduce GAUDI, a generative model capable of capturing the distribution of complex and realistic 3D scenes that can be rendered immersively from a moving camera. We tackle this challenging problem with a scalable yet powerful approach,…

Traditional 3D content creation tools empower users to bring their imagination to life by giving them direct control over a scene's geometry, appearance, motion, and camera path. Creating computer-generated videos, however, is a tedious…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shengqu Cai , Duygu Ceylan , Matheus Gadelha , Chun-Hao Paul Huang , Tuanfeng Yang Wang , Gordon Wetzstein

Recently, diffusion-based deep generative models (e.g., Stable Diffusion) have shown impressive results in text-to-image synthesis. However, current text-to-image models often require multiple passes of prompt engineering by humans in order…

Computation and Language · Computer Science 2023-11-14 Tingfeng Cao , Chengyu Wang , Bingyan Liu , Ziheng Wu , Jinhui Zhu , Jun Huang

We introduce a general framework for generating diverse visual content, including ambiguous images, panorama images, mesh textures, and Gaussian splat textures, by synchronizing multiple diffusion processes. We present exhaustive…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Jaihoon Kim , Juil Koo , Kyeongmin Yeo , Minhyuk Sung

The modeling and manipulation of 3D scenes captured from the real world are pivotal in various applications, attracting growing research interest. While previous works on editing have achieved interesting results through manipulating 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Guan Luo , Tian-Xing Xu , Ying-Tian Liu , Xiao-Xiong Fan , Fang-Lue Zhang , Song-Hai Zhang

3D scene reconstruction and rendering are core tasks in computer vision, with applications spanning industrial monitoring, robotics, and autonomous driving. Recent advances in 3D Gaussian Splatting (GS) and its variants have achieved…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Chi-Shiang Gau , Konstantinos D. Polyzos , Athanasios Bacharis , Saketh Madhuvarasu , Tara Javidi

We introduce Text2Cinemagraph, a fully automated method for creating cinemagraphs from text descriptions - an especially challenging task when prompts feature imaginary elements and artistic styles, given the complexity of interpreting the…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Aniruddha Mahapatra , Aliaksandr Siarohin , Hsin-Ying Lee , Sergey Tulyakov , Jun-Yan Zhu

Recent advancements in high-fidelity dynamic scene reconstruction have leveraged dynamic 3D Gaussians and 4D Gaussian Splatting for realistic scene representation. However, to make these methods viable for real-time applications such as…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Saqib Javed , Ahmad Jarrar Khan , Corentin Dumery , Chen Zhao , Mathieu Salzmann

This work introduces a new task of instance-incremental scene graph generation: Given a scene of the point cloud, representing it as a graph and automatically increasing novel instances. A graph denoting the object layout of the scene is…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Chao Qi , Jianqin Yin , Jinghang Xu , Pengxiang Ding

The recent Gaussian Splatting achieves high-quality and real-time novel-view synthesis of the 3D scenes. However, it is solely concentrated on the appearance and geometry modeling, while lacking in fine-grained object-level scene…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Mingqiao Ye , Martin Danelljan , Fisher Yu , Lei Ke

3D asset generation is getting massive amounts of attention, inspired by the recent success of text-guided 2D content creation. Existing text-to-3D methods use pretrained text-to-image diffusion models in an optimization problem or…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Lukas Höllein , Aljaž Božič , Norman Müller , David Novotny , Hung-Yu Tseng , Christian Richardt , Michael Zollhöfer , Matthias Nießner

Creating large-scale interactive 3D environments is essential for the development of Robotics and Embodied AI research. Current methods, including manual design, procedural generation, diffusion-based scene generation, and large language…

Computer Vision and Pattern Recognition · Computer Science 2024-11-18 Yian Wang , Xiaowen Qiu , Jiageng Liu , Zhehuan Chen , Jiting Cai , Yufei Wang , Tsun-Hsuan Wang , Zhou Xian , Chuang Gan

Recent advancements in photo-realistic novel view synthesis have been significantly driven by Gaussian Splatting (3DGS). Nevertheless, the explicit nature of 3DGS data entails considerable storage requirements, highlighting a pressing need…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Minye Wu , Tinne Tuytelaars

We propose a novel 3D deepfake generation framework based on 3D Gaussian Splatting that enables realistic, identity-preserving face swapping and reenactment in a fully controllable 3D space. Compared to conventional 2D deepfake approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Wending Liu , Siyun Liang , Huy H. Nguyen , Isao Echizen

Open-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understanding based on Gaussian splatting. However, these methods are…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yuxuan Xie , Xuan Yu , Changjian Jiang , Sitong Mao , Shunbo Zhou , Rui Fan , Rong Xiong , Yue Wang

Creating 4D fields of Gaussian Splatting from images or videos is a challenging task due to its under-constrained nature. While the optimization can draw photometric reference from the input videos or be regulated by generative models,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Quankai Gao , Qiangeng Xu , Zhe Cao , Ben Mildenhall , Wenchao Ma , Le Chen , Danhang Tang , Ulrich Neumann

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods that offer rich…

Generating 3D scenes is still a challenging task due to the lack of readily available scene data. Most existing methods only produce partial scenes and provide limited navigational freedom. We introduce a practical and scalable solution…

Graphics · Computer Science 2025-09-26 Zhaoyang Zhang , Yannick Hold-Geoffroy , Miloš Hašan , Ziwen Chen , Fujun Luan , Julie Dorsey , Yiwei Hu
‹ Prev 1 8 9 10 Next ›