English
Related papers

Related papers: OmniPSD: Layered PSD Generation with Diffusion Tra…

200 papers

Existing single image-to-3D creation methods typically involve a two-stage process, first generating multi-view images, and then using these images for 3D reconstruction. However, training these two stages separately leads to significant…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Hao Wen , Zehuan Huang , Yaohui Wang , Xinyuan Chen , Lu Sheng

Diffusion-based sparse-view CT (SVCT) imaging has achieved remarkable advancements in recent years, thanks to its more stable generative capability. However, recovering reliable image content and visually consistent textures is still a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Tianqi Wang , Wenchao Du , Hongyu Yang

Despite the groundbreaking success of diffusion models in generating high-fidelity images, their latent space remains relatively under-explored, even though it holds significant promise for enabling versatile and interpretable image editing…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Li Wang , Boyan Gao , Yanran Li , Zhao Wang , Xiaosong Yang , David A. Clifton , Jun Xiao

Diffusion-based text-to-image models ignited immense attention from the vision community, artists, and content creators. Broad adoption of these models is due to significant improvement in the quality of generations and efficient…

Computer Vision and Pattern Recognition · Computer Science 2023-10-19 Tianfu Wang , Menelaos Kanakis , Konrad Schindler , Luc Van Gool , Anton Obukhov

We present SPAD, a novel approach for creating consistent multi-view images from text prompts or single images. To enable multi-view generation, we repurpose a pretrained 2D diffusion model by extending its self-attention layers with…

Computer Vision and Pattern Recognition · Computer Science 2024-02-09 Yash Kant , Ziyi Wu , Michael Vasilkovsky , Guocheng Qian , Jian Ren , Riza Alp Guler , Bernard Ghanem , Sergey Tulyakov , Igor Gilitschenski , Aliaksandr Siarohin

Large-scale text-to-image generative models have been a revolutionary breakthrough in the evolution of generative AI, allowing us to synthesize diverse images that convey highly complex visual concepts. However, a pivotal challenge in…

Computer Vision and Pattern Recognition · Computer Science 2022-11-24 Narek Tumanyan , Michal Geyer , Shai Bagon , Tali Dekel

Automatic layout generation that can synthesize high-quality layouts is an important tool for graphic design in many applications. Though existing methods based on generative models such as Generative Adversarial Networks (GANs) and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Shang Chai , Liansheng Zhuang , Fengying Yan

Autoencoders empower state-of-the-art image and video generative models by compressing pixels into a latent space through visual tokenization. Although recent advances have alleviated the performance degradation of autoencoders under high…

Computer Vision and Pattern Recognition · Computer Science 2026-01-14 Dongxu Liu , Jiahui Zhu , Yuang Peng , Haomiao Tang , Yuwei Chen , Chunrui Han , Zheng Ge , Daxin Jiang , Mingxue Liao

We propose a novel PDE-driven corruption process for generative image synthesis based on advection-diffusion processes which generalizes existing PDE-based approaches. Our forward pass formulates image corruption via a physically motivated…

Graphics · Computer Science 2026-05-05 Grzegorz Gruszczynski , Jakub Meixner , Michal Jan Wlodarczyk , Przemyslaw Musialski

We introduce the Pyramid Diffusion Model (PDM), a novel architecture designed for ultra-high-resolution image synthesis. PDM utilizes a pyramid latent representation, providing a broader design space that enables more flexible, structured,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Jiajie Yang

Feed-forward 3D reconstruction for autonomous driving has advanced rapidly, yet existing methods struggle with the joint challenges of sparse, non-overlapping camera views and complex scene dynamics. We present UniSplat, a general…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Chen Shi , Shaoshuai Shi , Xiaoyang Lyu , Chunyang Liu , Kehua Sheng , Bo Zhang , Li Jiang

Score Distillation Sampling (SDS) has emerged as a prevalent technique for text-to-3D generation, enabling 3D content creation by distilling view-dependent information from text-to-2D guidance. However, they frequently exhibit shortcomings…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Zeyu Cai , Duotun Wang , Yixun Liang , Zhijing Shao , Ying-Cong Chen , Xiaohang Zhan , Zeyu Wang

Automatically generating high-quality real world 3D scenes is of enormous interest for applications such as virtual reality and robotics simulation. Towards this goal, we introduce NeuralField-LDM, a generative model capable of synthesizing…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Seung Wook Kim , Bradley Brown , Kangxue Yin , Karsten Kreis , Katja Schwarz , Daiqing Li , Robin Rombach , Antonio Torralba , Sanja Fidler

Generating complete 360-degree panoramas from narrow field of view images is ongoing research as omnidirectional RGB data is not readily available. Existing GAN-based approaches face some barriers to achieving higher quality output, and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Tianhao Wu , Chuanxia Zheng , Tat-Jen Cham

Probabilistic denoising diffusion models (DDMs) have set a new standard for 2D image generation. Extending DDMs for 3D content creation is an active field of research. Here, we propose TetraDiffusion, a diffusion model that operates on a…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Nikolai Kalischek , Torben Peters , Jan D. Wegner , Konrad Schindler

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

Radiance field representations have recently been explored in the latent space of VAEs that are commonly used by diffusion models. This direction offers efficient rendering and seamless integration with diffusion-based pipelines. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Or Hirschorn , Omer Sela , Inbar Huberman-Spiegelglas , Netalee Efrat , Eli Alshan , Ianir Ideses , Frederic Devernay , Yochai Zvik , Lior Fritz

Text-guided 3D face synthesis has achieved remarkable results by leveraging text-to-image (T2I) diffusion models. However, most existing works focus solely on the direct generation, ignoring the editing, restricting them from synthesizing…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Yunjie Wu , Yapeng Meng , Zhipeng Hu , Lincheng Li , Haoqian Wu , Kun Zhou , Weiwei Xu , Xin Yu

Diffusion models have recently set new benchmarks in Speech Enhancement (SE). However, most existing score-based models treat speech spectrograms merely as generic 2D images, applying uniform processing that ignores the intrinsic structural…

Sound · Computer Science 2026-02-03 Ke Xue , Rongfei Fan , Kai Li , Shanping Yu , Puning Zhao , Jianping An

Diffusion generative models transform noise into data by inverting a process that progressively adds noise to data samples. Inspired by concepts from the renormalization group in physics, which analyzes systems across different scales, we…

Machine Learning · Computer Science 2024-10-04 Mathis Gerdes , Max Welling , Miranda C. N. Cheng