English
Related papers

Related papers: Layout-and-Retouch: A Dual-stage Framework for Imp…

200 papers

The rapid development of diffusion models has triggered diverse applications. Identity-preserving text-to-image generation (ID-T2I) particularly has received significant attention due to its wide range of application scenarios like AI…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Weifeng Chen , Jiacheng Zhang , Jie Wu , Hefeng Wu , Xuefeng Xiao , Liang Lin

Benefited from image-text contrastive learning, pre-trained vision-language models, e.g., CLIP, allow to direct leverage texts as images (TaI) for parameter-efficient fine-tuning (PEFT). While CLIP is capable of making image features to be…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Chun-Mei Feng , Kai Yu , Xinxing Xu , Salman Khan , Rick Siow Mong Goh , Wangmeng Zuo , Yong Liu

Generative diffusion models can serve as a prior which ensures that solutions of image restoration systems adhere to the manifold of natural images. However, for restoring facial images, a personalized prior is necessary to accurately…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Pradyumna Chari , Sizhuo Ma , Daniil Ostashev , Achuta Kadambi , Gurunandan Krishnan , Jian Wang , Kfir Aberman

Despite recent advances in text-to-image (T2I) models, they often fail to faithfully render all elements of complex prompts, frequently omitting or misrepresenting specific objects and attributes. Test-time optimization has emerged as a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Mohammad Hossein Sameti , Amir M. Mansourian , Arash Marioriyad , Soheil Fadaee Oshyani , Mohammad Hossein Rohban , Mahdieh Soleymani Baghshah

Recent text-to-image (T2I) diffusion models have achieved remarkable progress in generating high-quality images given text-prompts as input. However, these models fail to convey appropriate spatial composition specified by a layout…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Jiayu Xiao , Henglei Lv , Liang Li , Shuhui Wang , Qingming Huang

Text-to-image (T2I) generation using multiple conditions enables fine-grained user control on the generated image. Yet, incorporating multi-condition inputs incurs substantial computation and communication overhead, due to additional…

Multimedia · Computer Science 2026-05-12 Yuxin Kong , Peng Yang , Chongbin Yi , Fan Wu , Feng Lyu

Text-guided 3D face synthesis has achieved remarkable results by leveraging text-to-image (T2I) diffusion models. However, most existing works focus solely on the direct generation, ignoring the editing, restricting them from synthesizing…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Yunjie Wu , Yapeng Meng , Zhipeng Hu , Lincheng Li , Haoqian Wu , Kun Zhou , Weiwei Xu , Xin Yu

The crux of text-to-image synthesis stems from the difficulty of preserving the cross-modality semantic consistency between the input text and the synthesized image. Typical methods, which seek to model the text-to-image mapping directly,…

Computer Vision and Pattern Recognition · Computer Science 2022-08-15 Jiadong Liang , Wenjie Pei , Feng Lu

Text-to-Image (T2I) models have transformed visual content creation, producing highly realistic images from natural language prompts. However, concerns persist around their potential to replicate and magnify existing societal biases. To…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Sedat Porikli , Vedat Porikli

Consistent text-to-image (T2I) generation seeks to produce identity-preserving images of the same subject across diverse scenes, yet it often fails due to a phenomenon called identity (ID) shift. Previous methods have tackled this issue,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Song Tang , Peihao Gong , Kunyu Li , Kai Guo , Boyu Wang , Mao Ye , Jianwei Zhang , Xiatian Zhu

Subject-driven text-to-image (T2I) generation aims to produce images that align with a given textual description, while preserving the visual identity from a referenced subject image. Despite its broad downstream applicability - ranging…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Aviv Slobodkin , Hagai Taitelbaum , Yonatan Bitton , Brian Gordon , Michal Sokolik , Nitzan Bitton Guetta , Almog Gueta , Royi Rassin , Dani Lischinski , Idan Szpektor

Text-to-image (T2I) diffusion models are popular for introducing image manipulation methods, such as editing, image fusion, inpainting, etc. At the same time, image-to-video (I2V) and text-to-video (T2V) models are also built on top of T2I…

A significant ``modality gap" exists between the abundance of text-only data and the increasing power of multimodal models. This work systematically investigates whether images generated on-the-fly by Text-to-Image (T2I) models can serve as…

Multimedia · Computer Science 2026-03-04 Yuesheng Huang , Peng Zhang , Xiaoxin Wu , Riliang Liu , Jiaqi Liang

Recent Text-to-Image (T2I) generation models such as Stable Diffusion and Imagen have made significant progress in generating high-resolution images based on text descriptions. However, many generated images still suffer from issues such as…

Large-scale text-to-image diffusion models have been a revolutionary milestone in the evolution of generative AI and multimodal technology, allowing wonderful image generation with natural-language text prompt. However, the issue of lacking…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xiang Gao , Jiaying Liu

Previous text-to-image synthesis algorithms typically use explicit textual instructions to generate/manipulate images accurately, but they have difficulty adapting to guidance in the form of coarsely matched texts. In this work, we attempt…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Mengyao Cui , Zhe Zhu , Shao-Ping Lu , Yulu Yang

Recent text-to-image (T2I) diffusion and flow-matching models can produce highly realistic images from natural language prompts. In practical scenarios, T2I systems are often run in a ``generate--then--select'' mode: many seeds are sampled…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Huanlei Guo , Hongxin Wei , Bingyi Jing

Modern vision models excel at general purpose downstream tasks. It is unclear, however, how they may be used for personalized vision tasks, which are both fine-grained and data-scarce. Recent works have successfully applied synthetic data…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Shobhita Sundaram , Julia Chae , Yonglong Tian , Sara Beery , Phillip Isola

Recent advancements in text-to-image (T2I) diffusion models have demonstrated remarkable capabilities in generating high-fidelity images. However, these models often struggle to faithfully render complex user prompts, particularly in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Linqing Wang , Ximing Xing , Yiji Cheng , Zhiyuan Zhao , Donghao Li , Tiankai Hang , Jiale Tao , Qixun Wang , Ruihuang Li , Comi Chen , Xin Li , Mingrui Wu , Xinchi Deng , Shuyang Gu , Chunyu Wang , Qinglin Lu

Personalized image retouching aims to adapt retouching style of individual users from reference examples, but existing methods often require user-specific fine-tuning or fail to generalize effectively. To address these challenges, we…

Graphics · Computer Science 2026-02-20 Temesgen Muruts Weldengus , Binnan Liu , Fei Kou , Youwei Lyu , Jinwei Chen , Qingnan Fan , Changqing Zou