English
Related papers

Related papers: ReSpace: Text-Driven Autoregressive 3D Indoor Scen…

200 papers

World building with 3D scene representations is increasingly important for content creation, simulation, and interactive experiences, yet real workflows are inherently iterative: creators must repeatedly extend an existing scene under user…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Zijian He , Renjie Liu , Yihao Wang , Weizhi Zhong , Huan Yuan , Kun Gai , Guangrun Wang , Guanbin Li

Versatile and adaptive semantic understanding would enable autonomous systems to comprehend and interact with their surroundings. Existing fixed-class models limit the adaptability of indoor mobile and assistive autonomous systems. In this…

Robotics · Computer Science 2024-03-06 Christina Kassab , Matias Mattamala , Lintong Zhang , Maurice Fallon

The demand for semantically rich 3D models of indoor scenes is rapidly growing, driven by applications in augmented reality, virtual reality, and robotics. However, creating them from sparse views remains a challenge due to geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Yijie Gao , Houqiang Zhong , Tianchi Zhu , Zhengxue Cheng , Qiang Hu , Li Song

This paper proposes a new method for simultaneous 3D reconstruction and semantic segmentation of indoor scenes. Unlike existing methods that require recording a video using a color camera and/or a depth camera, our method only needs a small…

Computer Vision and Pattern Recognition · Computer Science 2019-06-20 Jingyu Yang , Ji Xu , Kun Li , Yu-Kun Lai , Huanjing Yue , Jianzhi Lu , Hao Wu , Yebin Liu

Modern scene reconstruction methods are able to accurately recover 3D surfaces that are visible in one or more images. However, this leads to incomplete reconstructions, missing all occluded surfaces. While much progress has been made on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Sam Bahrami , Dylan Campbell

Synthesizing photo-realistic images and videos is at the heart of computer graphics and has been the focus of decades of research. Traditionally, synthetic images of a scene are generated using rendering algorithms such as rasterization or…

We present SemLayoutDiff, a unified model for synthesizing diverse 3D indoor scenes across multiple room types. The model introduces a scene layout representation combining a top-down semantic map and attributes for each object. Unlike…

Graphics · Computer Science 2025-09-09 Xiaohao Sun , Divyam Goel , Angel X. Chang

Designing 3D scenes is currently a creative task that requires significant expertise and effort in using complex 3D design interfaces. This effortful design process starts in stark contrast to the easiness with which people can use language…

Graphics · Computer Science 2017-03-02 Angel X. Chang , Mihail Eric , Manolis Savva , Christopher D. Manning

Furnishing and rendering indoor scenes has been a long-standing task for interior design, where artists create a conceptual design for the space, build a 3D model of the space, decorate, and then perform rendering. Although the task is…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Hong-Wing Pang , Yingshu Chen , Phuoc-Hieu Le , Binh-Son Hua , Duc Thanh Nguyen , Sai-Kit Yeung

Despite recent progress in 3D-LLMs, they remain limited in accurately grounding language to visual and spatial elements in 3D environments. This limitation stems in part from training data that focuses on language reasoning rather than…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Yue Zhang , Zun Wang , Han Lin , Jialu Li , Jianing Yang , Yonatan Bitton , Idan Szpektor , Mohit Bansal

Synthetic data has been a critical tool for training scene text detection and recognition models. On the one hand, synthetic word images have proven to be a successful substitute for real images in training scene text recognizers. On the…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Shangbang Long , Cong Yao

We present BlenderFusion, a generative visual compositing framework that synthesizes new scenes by recomposing objects, camera, and background. It follows a layering-editing-compositing pipeline: (i) segmenting and converting visual inputs…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jiacheng Chen , Ramin Mehran , Xuhui Jia , Saining Xie , Sanghyun Woo

Traditionally, 3D scene synthesis requires expert knowledge and significant manual effort. Automating this process could greatly benefit fields such as architectural design, robotics simulation, virtual reality, and gaming. Recent…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Rui Huang , Guangyao Zhai , Zuria Bauer , Marc Pollefeys , Federico Tombari , Leonidas Guibas , Gao Huang , Francis Engelmann

Most deep learning approaches to comprehensive semantic modeling of 3D indoor spaces require costly dense annotations in the 3D domain. In this work, we explore a central 3D scene modeling task, namely, semantic scene reconstruction without…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Junwen Huang , Alexey Artemov , Yujin Chen , Shuaifeng Zhi , Kai Xu , Matthias Nießner

Text-to-image models are showcasing the impressive ability to create high-quality and diverse generative images. Nevertheless, the transition from freehand sketches to complex scene images remains challenging using diffusion models. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Tianyu Zhang , Xiaoxuan Xie , Xusheng Du , Haoran Xie

Image generation models trained on large datasets can synthesize high-quality images but often produce spatially inconsistent and distorted images due to limited information about the underlying structures and spatial layouts. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Hyundo Lee , Suhyung Choi , Inwoo Hwang , Byoung-Tak Zhang

Scene text synthesis involves rendering specified texts onto arbitrary images. Current methods typically formulate this task in an end-to-end manner but lack effective character-level guidance during training. Besides, their text encoders,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yibin Wang , Weizhong Zhang , Honghui Xu , Cheng Jin

We are witnessing significant breakthroughs in the technology for generating 3D objects from text. Existing approaches either leverage large text-to-image models to optimize a 3D representation or train 3D generators on object-centric…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Qihang Zhang , Chaoyang Wang , Aliaksandr Siarohin , Peiye Zhuang , Yinghao Xu , Ceyuan Yang , Dahua Lin , Bolei Zhou , Sergey Tulyakov , Hsin-Ying Lee

Visual Language Models (VLMs) have increasingly become the main paradigm for understanding indoor scenes, but they still struggle with metric and spatial reasoning. Current approaches rely on end-to-end video understanding or large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Fernando Ropero , Erkin Turkoz , Daniel Matos , Junqing Du , Antonio Ruiz , Yanfeng Zhang , Lu Liu , Mingwei Sun , Yongliang Wang

We present a novel approach for 3D indoor scene reconstruction that combines 3D Gaussian Splatting (3DGS) with mesh representations. We use meshes for the room layout of the indoor scene, such as walls, ceilings, and floors, while employing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Jiyeop Kim , Jongwoo Lim
‹ Prev 1 4 5 6 7 8 10 Next ›