English
Related papers

Related papers: Maestro: Self-Improving Text-to-Image Generation v…

200 papers

This paper presents a novel approach to enhance image-to-image generation by leveraging the multimodal capabilities of the Large Language and Vision Assistant (LLaVA). We propose a framework where LLaVA analyzes input images and generates…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Zhicheng Ding , Panfeng Li , Qikai Yang , Siyang Li

Text-to-image (T2I) models have substantially improved image fidelity and prompt adherence, yet their creativity remains constrained by reliance on discrete natural language prompts. When presented with fuzzy prompts such as ``a creative…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Ruixiao Shi , Fu Feng , Yucheng Xie , Xu Yang , Jing Wang , Xin Geng

Despite the impressive advances in text-to-image models, they often struggle to effectively compose complex scenes with multiple objects, displaying various attributes and relationships. To address this challenge, we present…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Kaiyi Huang , Chengqi Duan , Kaiyue Sun , Enze Xie , Zhenguo Li , Xihui Liu

Despite advancements in text-to-image generation (T2I), prior methods often face text-image misalignment problems such as relation confusion in generated images. Existing solutions involve cross-attention manipulation for better…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Leigang Qu , Wenjie Wang , Yongqi Li , Hanwang Zhang , Liqiang Nie , Tat-Seng Chua

Recently, text-to-image (T2I) synthesis has undergone significant advancements, particularly with the emergence of Large Language Models (LLM) and their enhancement in Large Vision Models (LVM), greatly enhancing the instruction-following…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Weijin Cheng , Jianzhi Liu , Jiawen Deng , Fuji Ren

Recent advancements in text-to-image (T2I) diffusion models have demonstrated remarkable capabilities in generating high-fidelity images. However, these models often struggle to faithfully render complex user prompts, particularly in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Linqing Wang , Ximing Xing , Yiji Cheng , Zhiyuan Zhao , Donghao Li , Tiankai Hang , Jiale Tao , Qixun Wang , Ruihuang Li , Comi Chen , Xin Li , Mingrui Wu , Xinchi Deng , Shuyang Gu , Chunyu Wang , Qinglin Lu

Well-designed prompts can guide text-to-image models to generate amazing images. However, the performant prompts are often model-specific and misaligned with user input. Instead of laborious human engineering, we propose prompt adaptation,…

Computation and Language · Computer Science 2024-01-01 Yaru Hao , Zewen Chi , Li Dong , Furu Wei

Text-to-image (T2I) systems enable rapid generation of high-fidelity imagery but are misaligned with how visual ideas develop. T2I systems generate outputs that make implicit visual decisions on behalf of the user, often introduce…

Human-Computer Interaction · Computer Science 2026-04-16 Zoe De Simone , Angie Boggust , Fredo Durand , Ashia Wilson , Arvind Satyanarayan

Recent advancements in text-to-image generation have revolutionized numerous fields, including art and cinema, by automating the generation of high-quality, context-aware images and video. However, the utility of these technologies is often…

The revolution of artificial intelligence content generation has been rapidly accelerated with the booming text-to-image (T2I) diffusion models. Within just two years of development, it was unprecedentedly of high-quality, diversity, and…

Artificial Intelligence · Computer Science 2023-10-16 Zeqiang Lai , Xizhou Zhu , Jifeng Dai , Yu Qiao , Wenhai Wang

Text-to-image generative models have achieved remarkable visual quality but still struggle with compositionality$-$accurately capturing object relationships, attribute bindings, and fine-grained details in prompts. A key limitation is that…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Arman Zarei , Jiacheng Pan , Matthew Gwilliam , Soheil Feizi , Zhenheng Yang

Recent advancements in text-to-image (T2I) generation have enabled models to produce high-quality images from textual descriptions. However, these models often struggle with complex instructions involving multiple objects, attributes, and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Yucheng Zhou , Jiahao Yuan , Qianning Wang

Generative text-to-image models have gained great popularity among the public for their powerful capability to generate high-quality images based on natural language prompts. However, developing effective prompts for desired images can be…

Artificial Intelligence · Computer Science 2023-11-02 Yingchaojie Feng , Xingbo Wang , Kam Kwai Wong , Sijia Wang , Yuhong Lu , Minfeng Zhu , Baicheng Wang , Wei Chen

Despite their impressive capabilities, diffusion-based text-to-image (T2I) models can lack faithfulness to the text prompt, where generated images may not contain all the mentioned objects, attributes or relations. To alleviate these…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Shyamgopal Karthik , Karsten Roth , Massimiliano Mancini , Zeynep Akata

Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, particularly those involving multiple subjects with distinct…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Lifeng Chen , Jiner Wang , Zihao Pan , Beier Zhu , Xiaofeng Yang , Chi Zhang

Text-to-image generation model is able to generate images across a diverse range of subjects and styles based on a single prompt. Recent works have proposed a variety of interaction methods that help users understand the capabilities of…

Human-Computer Interaction · Computer Science 2023-07-19 Seungho Baek , Hyerin Im , Jiseung Ryu , Juhyeong Park , Takyeon Lee

In the text-to-image generation field, recent remarkable progress in Stable Diffusion makes it possible to generate rich kinds of novel photorealistic images. However, current models still face misalignment issues (e.g., problematic spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Leigang Qu , Shengqiong Wu , Hao Fei , Liqiang Nie , Tat-Seng Chua

Text-to-image (T2I) diffusion models such as SDXL and FLUX have achieved impressive photorealism, yet small-scale distortions remain pervasive in limbs, face, text and so on. Existing refinement approaches either perform costly iterative…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Shaocheng Shen , Jianfeng Liang , Chunlei Cai , Cong Geng , Huiyu Duan , Xiaoyun Zhang , Qiang Hu , Guangtao Zhai

Text-to-image generation (T2I) refers to the text-guided generation of high-quality images. In the past few years, T2I has attracted widespread attention and numerous works have emerged. In this survey, we comprehensively review 141 works…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Pengfei Yang , Ngai-Man Cheung , Xinda Ma

Recently, the strong latent Diffusion Probabilistic Model (DPM) has been applied to high-quality Text-to-Image (T2I) generation (e.g., Stable Diffusion), by injecting the encoded target text prompt into the gradually denoised diffusion…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Mingyang Yi , Aoxue Li , Yi Xin , Zhenguo Li