English
Related papers

Related papers: Proactive Agents for Multi-Turn Text-to-Image Gene…

200 papers

With the rapid advancement of large multimodal models (LMMs), recent text-to-image (T2I) models can generate high-quality images and demonstrate great alignment to short prompts. However, they still struggle to effectively understand and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Juntong Wang , Huiyu Duan , Jiarui Wang , Ziheng Jia , Guangtao Zhai , Xiongkuo Min

Recent text-to-image (T2I) diffusion models have achieved remarkable progress in generating high-quality images given text-prompts as input. However, these models fail to convey appropriate spatial composition specified by a layout…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Jiayu Xiao , Henglei Lv , Liang Li , Shuhui Wang , Qingming Huang

Well-designed prompts can guide text-to-image models to generate amazing images. However, the performant prompts are often model-specific and misaligned with user input. Instead of laborious human engineering, we propose prompt adaptation,…

Computation and Language · Computer Science 2024-01-01 Yaru Hao , Zewen Chi , Li Dong , Furu Wei

Text-to-image models are enabling efficient design space exploration, rapidly generating images from text prompts. However, many generative AI tools are imperfect for product design applications as they are not built for the goals and…

Human-Computer Interaction · Computer Science 2025-01-22 Leah Chong , I-Ping Lo , Jude Rayan , Steven Dow , Faez Ahmed , Ioanna Lykourentzou

The emergence of generative AI (GenAI) models, including large language models and text-to-image models, has significantly advanced the synergy between humans and AI with not only their outstanding capability but more importantly, the…

Human-Computer Interaction · Computer Science 2025-03-05 Leixian Shen , Haotian Li , Yifang Wang , Xing Xie , Huamin Qu

Recent advancements in text-to-image (T2I) generative models have shown remarkable capabilities in producing diverse and imaginative visuals based on text prompts. Despite the advancement, these diffusion models sometimes struggle to…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Xiaohui Chen , Yongfei Liu , Yingxiang Yang , Jianbo Yuan , Quanzeng You , Li-Ping Liu , Hongxia Yang

Recent developments in large language models (LLM) and generative AI have unleashed the astonishing capabilities of text-to-image generation systems to synthesize high-quality images that are faithful to a given reference text, known as a…

Human-Computer Interaction · Computer Science 2023-03-17 Yutong Xie , Zhaoying Pan , Jinge Ma , Luo Jie , Qiaozhu Mei

Text-to-Image (T2I) diffusion models are widely recognized for their ability to generate high-quality and diverse images based on text prompts. However, despite recent advances, these models are still prone to generating unsafe images…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Jiangweizhi Peng , Zhiwei Tang , Gaowen Liu , Charles Fleming , Mingyi Hong

Recently, prompt learning has emerged as the state-of-the-art (SOTA) for fair text-to-image (T2I) generation. Specifically, this approach leverages readily available reference images to learn inclusive prompts for each target Sensitive…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Christopher T. H Teo , Milad Abdollahzadeh , Xinda Ma , Ngai-man Cheung

The misuse of generative AI in online disinformation campaigns highlights the urgent need for transparent and explainable detection systems. In this work, we investigate how detectors for AI-generated images can be more effective in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Silvia Poletti , Justin Ilyes , Marcel Hasenbalg , David Fischinger , Martin Boyer

Generating images from textual descriptions has recently attracted a lot of interest. While current models can generate photo-realistic images of individual objects such as birds and human faces, synthesising images with multiple objects is…

Computer Vision and Pattern Recognition · Computer Science 2020-10-29 Stanislav Frolov , Shailza Jolly , Jörn Hees , Andreas Dengel

Recent progress in text-to-image (T2I) diffusion models (DMs) has enabled high-quality visual synthesis from diverse textual prompts. Yet, most existing T2I DMs, even those equipped with large language model (LLM)-based text encoders,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Siqi Kou , Jiachun Jin , Zetong Zhou , Ye Ma , Yugang Wang , Quan Chen , Peng Jiang , Xiao Yang , Jun Zhu , Kai Yu , Zhijie Deng

Neural image classifiers are known to undergo severe performance degradation when exposed to inputs that are sampled from environmental conditions that differ from their training data. Given the recent progress in Text-to-Image (T2I)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Jianhao Yuan , Francesco Pinto , Adam Davies , Philip Torr

Recent text-to-image (T2I) diffusion and flow-matching models can produce highly realistic images from natural language prompts. In practical scenarios, T2I systems are often run in a ``generate--then--select'' mode: many seeds are sampled…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Huanlei Guo , Hongxin Wei , Bingyi Jing

Text-to-image (T2I) models based on diffusion and transformer architectures advance rapidly. They are often pretrained on large corpora, and openly shared on a model platform, such as HuggingFace. Users can then build up AI applications,…

Machine Learning · Computer Science 2025-08-18 Basile Lewandowski , Robert Birke , Lydia Y. Chen

Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images according to prompt alignment and perceptual quality. Existing reward models are commonly…

Artificial Intelligence · Computer Science 2026-05-22 Kuei-Chun Kao , Daixuan Huo , Yuanhao Ban , Cho-Jui Hsieh

Text-to-image models take a sentence (i.e., prompt) and generate images associated with this input prompt. These models have created award wining-art, videos, and even synthetic datasets. However, text-to-image (T2I) models can generate…

Computation and Language · Computer Science 2023-06-12 Alexander Lin , Lucas Monteiro Paes , Sree Harsha Tanneru , Suraj Srinivas , Himabindu Lakkaraju

Ethical issues around text-to-image (T2I) models demand a comprehensive control over the generative content. Existing techniques addressing these issues for responsible T2I models aim for the generated content to be fair and safe…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Basim Azam , Naveed Akhtar

Text-to-image (T2I) models have advanced considerably in generating high-quality images from textual descriptions. However, their ability to associate colors with concepts remains largely constrained to explicit color names or codes, while…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Chenxi Ruan , Yihan Hou , Yu Xiao , Guosheng Hu , Wei Zeng

Recent advancements in text-to-image (T2I) generation models have transformed the field. However, challenges persist in generating images that reflect demanding textual descriptions, especially for fine-grained details and unusual…

Multimedia · Computer Science 2025-02-21 Ran Li , Xiaomeng Jin , Heng ji