English
Related papers

Related papers: REAL: Realism Evaluation of Text-to-Image Generati…

200 papers

With the rapid evolution of the Text-to-Image (T2I) model in recent years, their unsatisfactory generation result has become a challenge. However, uniformly refining AI-Generated Images (AIGIs) of different qualities not only limited…

Computer Vision and Pattern Recognition · Computer Science 2024-01-03 Chunyi Li , Haoning Wu , Zicheng Zhang , Hongkun Hao , Kaiwei Zhang , Lei Bai , Xiaohong Liu , Xiongkuo Min , Weisi Lin , Guangtao Zhai

Recent advances in text-driven image editing have been significant, yet the task of accurately evaluating these edited images continues to pose a considerable challenge. Different from the assessment of text-driven image generation,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Bowen Qu , Shangkun Sun , Xiaoyu Liang , Wei Gao

Text-to-image (T2I) generative models can create vivid, realistic images from textual descriptions. As these models proliferate, they expose new concerns about their ability to represent diverse demographic groups, propagate stereotypes,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Sangwon Jung , Alex Oesterling , Claudio Mayrink Verdun , Sajani Vithana , Taesup Moon , Flavio P. Calmon

Text-to-image (T2I) models today are capable of producing photorealistic, instruction-following images, yet they still frequently fail on prompts that require implicit world knowledge. Existing evaluation protocols either emphasize…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Tianyang Han , Junhao Su , Junjie Hu , Peizhen Yang , Hengyu Shi , Junfeng Luo , Jialin Gao

Fine-tuning text-to-image models with reward functions trained on human feedback data has proven effective for aligning model behavior with human intent. However, excessive optimization with such reward models, which serve as mere proxy…

Machine Learning · Computer Science 2024-04-03 Kyuyoung Kim , Jongheon Jeong , Minyong An , Mohammad Ghavamzadeh , Krishnamurthy Dvijotham , Jinwoo Shin , Kimin Lee

Despite advances in generation quality, current text-to-image (T2I) models often lack diversity, generating homogeneous outputs. This work introduces a framework to address the need for robust diversity evaluation in T2I models. Our…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Isabela Albuquerque , Ira Ktena , Olivia Wiles , Ivana Kajić , Amal Rannen-Triki , Cristina Vasconcelos , Aida Nematzadeh

In the text-to-image generation field, recent remarkable progress in Stable Diffusion makes it possible to generate rich kinds of novel photorealistic images. However, current models still face misalignment issues (e.g., problematic spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Leigang Qu , Shengqiong Wu , Hao Fei , Liqiang Nie , Tat-Seng Chua

Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion and video editing. A specialized area within generative…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Jiaxin Cheng , Zixu Zhao , Tong He , Tianjun Xiao , Yicong Zhou , Zheng Zhang

This paper explores the feasibility of using text-to-image models in a zero-shot setup to generate images for taxonomy concepts. While text-based methods for taxonomy enrichment are well-established, the potential of the visual dimension…

Computation and Language · Computer Science 2025-03-14 Viktor Moskvoretskii , Alina Lobanova , Ekaterina Neminova , Chris Biemann , Alexander Panchenko , Irina Nikishina

Evaluating the alignment between textual prompts and generated images is critical for ensuring the reliability and usability of text-to-image (T2I) models. However, most existing evaluation methods rely on coarse-grained metrics or static…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Fulin Shi , Wenyi Xiao , Bin Chen , Liang Din , Leilei Gan

Text-to-image (T2I) diffusion models have become prominent tools for generating high-fidelity images from text prompts. However, when trained on unfiltered internet data, these models can produce unsafe, incorrect, or stylistically…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Rohit Jena , Ali Taghibakhshi , Sahil Jain , Gerald Shen , Nima Tajbakhsh , Arash Vahdat

Continual post-training adapts a single text-to-image diffusion model to learn new tasks without incurring the cost of separate models, but naive post-training causes forgetting of pretrained knowledge and undermines zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Zhehao Huang , Yuhang Liu , Yixin Lou , Zhengbao He , Mingzhen He , Wenxing Zhou , Tao Li , Kehan Li , Zeyi Huang , Xiaolin Huang

The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate the T2I progress. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Jiahui Chen , Candace Ross , Reyhane Askari-Hemmat , Koustuv Sinha , Melissa Hall , Michal Drozdzal , Adriana Romero-Soriano

In this paper, we investigate the potential of image-to-image translation (I2I) techniques for transferring realism to 3D-rendered facial images in the context of Face Recognition (FR) systems. The primary motivation for using 3D-rendered…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Parsa Rahimi , Behrooz Razeghi , Sebastien Marcel

Text-to-image generation and text-guided image manipulation have received considerable attention in the field of image generation tasks. However, the mainstream evaluation methods for these tasks have difficulty in evaluating whether all…

Computer Vision and Pattern Recognition · Computer Science 2024-11-18 Mizuki Miyamoto , Ryugo Morita , Jinjia Zhou

User prompts for generative AI models are often underspecified, leading to a misalignment between the user intent and models' understanding. As a result, users commonly have to painstakingly refine their prompts. We study this alignment…

Artificial Intelligence · Computer Science 2025-10-27 Meera Hahn , Wenjun Zeng , Nithish Kannen , Rich Galt , Kartikeya Badola , Been Kim , Zi Wang

Recent text-to-image (T2I) diffusion and flow-matching models can produce highly realistic images from natural language prompts. In practical scenarios, T2I systems are often run in a ``generate--then--select'' mode: many seeds are sampled…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Huanlei Guo , Hongxin Wei , Bingyi Jing

Multimodal text-to-image generation remains constrained by the difficulty of maintaining semantic alignment and professional-level detail across diverse visual domains. We propose a multi-agent reinforcement learning framework that…

Artificial Intelligence · Computer Science 2025-10-14 Jiabao Shi , Minfeng Qi , Lefeng Zhang , Di Wang , Yingjie Zhao , Ziying Li , Yalong Xing , Ningran Li

Current text-to-image generative models struggle to accurately represent object states (e.g., "a table without a bottle," "an empty tumbler"). In this work, we first design a fully-automatic pipeline to generate high-quality synthetic data…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Tianle Chen , Chaitanya Chakka , Deepti Ghadiyaram

The unprecedented photorealistic results achieved by recent text-to-image generative systems and their increasing use as plug-and-play content creation solutions make it crucial to understand their potential biases. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Melissa Hall , Candace Ross , Adina Williams , Nicolas Carion , Michal Drozdzal , Adriana Romero Soriano