English
Related papers

Related papers: HierEdit: Region-Aware Hierarchical Diffusion for …

200 papers

Text-to-image diffusion models can generate diverse, high-fidelity images based on user-provided text prompts. Recent research has extended these models to support text-guided image editing. While text guidance is an intuitive editing…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Jooyoung Choi , Yunjey Choi , Yunji Kim , Junho Kim , Sungroh Yoon

Colour is one of the most perceptually salient yet least controllable attributes in image generation. Although recent diffusion models can modify object colours from user instructions, their results often deviate from the intended hue,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Yuqi Yang , Dongliang Chang , Yijia Ling , Ruoyi Du , Zhanyu Ma

Recent advances in diffusion models have enabled high-quality generation and manipulation of images guided by texts, as well as concept learning from images. However, naive applications of existing methods to editing tasks that require…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Xudong Liu , Zikun Chen , Ruowei Jiang , Ziyi Wu , Kejia Yin , Han Zhao , Parham Aarabi , Igor Gilitschenski

Despite recent advances in large-scale text-to-image generative models, manipulating real images with these models remains a challenging problem. The main limitations of existing editing methods are that they either fail to perform with…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Vadim Titov , Madina Khalmatova , Alexandra Ivanova , Dmitry Vetrov , Aibek Alanov

Diffusion models have revolutionized text-to-image generation, but their real-world applications are hampered by the extensive time needed for hundreds of diffusion steps. Although progressive distillation has been proposed to speed up…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Yifan Zhang , Bryan Hooi

Users often possess a clear visual intent but struggle to articulate it precisely in language. This intention-expression gap makes aligning generated images with latent visual preferences a fundamental challenge in text-to-image diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Wenxi Wang , Hongbin Liu , Mingqian Li , Junyan Yuan , Junqi Zhang

Diffusion models have attained impressive visual quality for image synthesis. However, how to interpret and manipulate the latent space of diffusion models has not been extensively explored. Prior work diffusion autoencoders encode the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-26 Zeyu Lu , Chengyue Wu , Xinyuan Chen , Yaohui Wang , Lei Bai , Yu Qiao , Xihui Liu

Instruction-based image editing has made a great process in using natural human language to manipulate the visual content of images. However, existing models are limited by the quality of the dataset and cannot accurately localize editing…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Tiancheng Li , Jinxiu Liu , Huajun Chen , Qi Liu

Nonprehensile manipulation, such as pushing objects across cluttered environments, presents a challenging control problem due to complex contact dynamics and long-horizon planning requirements. In this work, we propose HeRD, a hierarchical…

Robotics · Computer Science 2025-12-12 Steven Caro , Stephen L. Smith

Diffusion models have emerged as the leading approach for image synthesis, demonstrating exceptional photorealism and diversity. However, training diffusion models at high resolutions remains computationally prohibitive, and existing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Tobias Vontobel , Seyedmorteza Sadat , Farnood Salehi , Romann M. Weber

Conditional diffusion models have demonstrated impressive performance in image manipulation tasks. The general pipeline involves adding noise to the image and then denoising it. However, this method faces a trade-off problem: adding too…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Luozhou Wang , Shuai Yang , Shu Liu , Ying-cong Chen

Diffusion-based image editing has made semantic level image manipulation easy for general users, but it also enables realistic local forgeries that are hard to localize. Existing benchmarks mainly focus on the binary detection of generated…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Rui Zhang , Hongxia Wang , Hangqing Liu , Yang Zhou , Qiang Zeng

Image restoration is a classic low-level problem aimed at recovering high-quality images from low-quality images with various degradations such as blur, noise, rain, haze, etc. However, due to the inherent complexity and non-uniqueness of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Yuhong Zhang , Hengsheng Zhang , Xinning Chai , Zhengxue Cheng , Rong Xie , Li Song , Wenjun Zhang

Diffusion models excel at producing high-quality images; however, scaling to higher resolutions, such as 4K, often results in over-smoothed content, structural distortions, and repetitive patterns. To this end, we introduce ResMaster, a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Shuwei Shi , Wenbo Li , Yuechen Zhang , Jingwen He , Biao Gong , Yinqiang Zheng

Deep image restoration models aim to learn a mapping from degraded image space to natural image space. However, they face several critical challenges: removing degradation, generating realistic details, and ensuring pixel-level consistency.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Xinqi Lin , Fanghua Yu , Jinfan Hu , Zhiyuan You , Wu Shi , Jimmy S. Ren , Jinjin Gu , Chao Dong

Open-domain 3D object synthesis has been lagging behind image synthesis due to limited data and higher computational complexity. To bridge this gap, recent works have investigated multi-view diffusion but often fall short in either 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Hansheng Chen , Ruoxi Shi , Yulin Liu , Bokui Shen , Jiayuan Gu , Gordon Wetzstein , Hao Su , Leonidas Guibas

Text-to-image (T2I) diffusion models, with their impressive generative capabilities, have been adopted for image editing tasks, demonstrating remarkable efficacy. However, due to attention leakage and collision between the cross-attention…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Xingxi Yin , Zhi Li , Jingfeng Zhang , Chenglin Li , Yin Zhang

3D content creation from a single image is a long-standing yet highly desirable task. Recent advances introduce 2D diffusion priors, yielding reasonable results. However, existing methods are not hyper-realistic enough for post-generation…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Tong Wu , Zhibing Li , Shuai Yang , Pan Zhang , Xinggang Pan , Jiaqi Wang , Dahua Lin , Ziwei Liu

Ultra-high-resolution text-to-image generation is increasingly vital for applications requiring fine-grained textures and global structural fidelity, yet state-of-the-art text-to-image diffusion models such as FLUX and SD3 remain confined…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Yuyao Zhang , Yu-Wing Tai

Generative models are increasingly used to augment medical imaging datasets for fairer AI. Yet a key assumption often goes unexamined: that generators themselves produce equally high-quality images across demographic groups. Models trained…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Mahmoud Ibrahim , Bart Elen , Chang Sun , Gokhan Ertaylan , Michel Dumontier