English
Related papers

Related papers: CymbaDiff: Structured Spatial Diffusion for Sketch…

200 papers

Existing text-based 3D generation methods generate attractive results but lack detailed geometry control. Sketches, known for their conciseness and expressiveness, have contributed to intuitive 3D modeling but are confined to producing…

Graphics · Computer Science 2024-05-15 Feng-Lin Liu , Hongbo Fu , Yu-Kun Lai , Lin Gao

This work addresses a gap in semantic scene completion (SSC) data by creating a novel outdoor data set with accurate and complete dynamic scenes. Our data set is formed from randomly sampled views of the world at each time step, which…

Computer Vision and Pattern Recognition · Computer Science 2022-07-01 Joey Wilson , Jingyu Song , Yuewei Fu , Arthur Zhang , Andrew Capodieci , Paramsothy Jayakumar , Kira Barton , Maani Ghaffari

Synthesizing realistic 3D indoor scenes remains challenging due to data scarcity and the difficulty of simultaneously enforcing global architectural constraints and local semantic consistency. Existing approaches often overlook structural…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Yingrui Wu , Youkang Kong , Mingyang Zhao , Weize Quan , Dong-Ming Yan , Yang Liu

We present SceneFactor, a diffusion-based approach for large-scale 3D scene generation that enables controllable generation and effortless editing. SceneFactor enables text-guided 3D scene synthesis through our factored diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Alexey Bokhovkin , Quan Meng , Shubham Tulsiani , Angela Dai

Street-view imagery (SVI) is widely used to quantify key indicators of urban environment, such as green- ery, sky, or road view indices. However, existing studies largely focus on measuring current streetscapes and rarely support the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yuzhou Chen , Yuebing Liang , Lingqian Hu , Kailai Sun , Qingqi Song , Chang Zhao , Shenhao Wang

Semantic understanding of scenes in three-dimensional space (3D) is a quintessential part of robotics oriented applications such as autonomous driving as it provides geometric cues such as size, orientation and true distance of separation…

Computer Vision and Pattern Recognition · Computer Science 2019-11-01 Kartik Srivastava , Akash Kumar Singh , Guruprasad M. Hegde

Diffusion-based policies have shown remarkable capability in executing complex robotic manipulation tasks but lack explicit characterization of geometry and semantics, which often limits their ability to generalize to unseen objects and…

Robotics · Computer Science 2024-10-24 Yixuan Wang , Guang Yin , Binghao Huang , Tarik Kelestemur , Jiuguang Wang , Yunzhu Li

In the medical domain, acquiring large datasets is challenging due to both accessibility issues and stringent privacy regulations. Consequently, data availability and privacy protection are major obstacles to applying machine learning in…

Image and Video Processing · Electrical Eng. & Systems 2025-07-02 Wenwu Tang , Khaled Seyam , Bin Yang

We introduce a novel state-space architecture for diffusion models, effectively harnessing spatial and frequency information to enhance the inductive bias towards local features in input images for image generation tasks. While state-space…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Hao Phung , Quan Dao , Trung Dao , Hoang Phan , Dimitris Metaxas , Anh Tran

We present LT3SD, a novel latent diffusion model for large-scale 3D scene generation. Recent advances in diffusion models have shown impressive results in 3D object generation, but are limited in spatial extent and quality when extended to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Quan Meng , Lei Li , Matthias Nießner , Angela Dai

Perception systems play a crucial role in autonomous driving, incorporating multiple sensors and corresponding computer vision algorithms. 3D LiDAR sensors are widely used to capture sparse point clouds of the vehicle's surroundings.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Helin Cao , Sven Behnke

Training autonomous driving and navigation systems requires large and diverse point cloud datasets that capture complex edge case scenarios from various dynamic urban settings. Acquiring such diverse scenarios from real-world point cloud…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Suchetan G. Uppur , Hemant Kumar , Vaibhav Kumar

Designing complex 3D scenes has been a tedious, manual process requiring domain expertise. Emerging text-to-3D generative models show great promise for making this task more intuitive, but existing approaches are limited to object-level…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Ryan Po , Gordon Wetzstein

Synthetic datasets are widely used for training urban scene recognition models, but even highly realistic renderings show a noticeable gap to real imagery. This gap is particularly pronounced when adapting to a specific target domain, such…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Denis Zavadski , Damjan Kalšan , Tim Küchler , Haebom Lee , Stefan Roth , Carsten Rother

Surgical scene segmentation is essential for enhancing surgical precision, yet it is frequently compromised by the scarcity and imbalance of available data. To address these challenges, semantic image synthesis methods based on generative…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Yihang Zhou , Rebecca Towning , Zaid Awad , Stamatia Giannarou

The completion, extension, and generation of 3D semantic scenes are an interrelated set of capabilities that are useful for robotic navigation and exploration. Existing approaches seek to decouple these problems and solve them one-off.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Xujia Zhang , Brendan Crowe , Christoffer Heckman

Text-to-image models are showcasing the impressive ability to create high-quality and diverse generative images. Nevertheless, the transition from freehand sketches to complex scene images remains challenging using diffusion models. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Tianyu Zhang , Xiaoxuan Xie , Xusheng Du , Haoran Xie

3D conducting motion generation aims to synthesize fine-grained conductor motions from music, with broad potential in music education, virtual performance, digital human animation, and human-AI co-creation. However, this task remains…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Tianzhi Jia , Kaixing Yang , Xiaole Yang , Xulong Tang , Ke Qiu , Shikui Wei , Yao Zhao

Text-to-image models can generate visually appealing images from text descriptions. Efforts have been devoted to improving model controls with prompt tuning and spatial conditioning. However, our formative study highlights the challenges…

Human-Computer Interaction · Computer Science 2025-02-12 Haichuan Lin , Yilin Ye , Jiazhi Xia , Wei Zeng

Designing stylized cinemagraphs is challenging due to the difficulty in customizing complex and expressive flow elements. To achieve intuitive and detailed control of the generated cinemagraphs, sketches provide a feasible solution to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Hao Jin , Hengyuan Chang , Xiaoxuan Xie , Zhengyang Wang , Xusheng Du , Shaojun Hu , Haoran Xie