English
Related papers

Related papers: GenEval 2: Addressing Benchmark Drift in Text-to-I…

200 papers

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

High-quality and open datasets remain a major bottleneck for text-to-image (T2I) fine-tuning. Despite rapid progress in model architectures and training pipelines, most publicly available fine-tuning datasets suffer from low resolution,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Xu Ma , Yitian Zhang , Qihua Dong , Yun Fu

While recent advancements in generative modeling have significantly improved text-image alignment, some residual misalignment between text and image representations still remains. Some approaches address this issue by fine-tuning models in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Jaa-Yeon Lee , Byunghee Cha , Jeongsol Kim , Jong Chul Ye

The recent advancement of large and powerful models with Text-to-Image (T2I) generation abilities -- such as OpenAI's DALLE-3 and Google's Gemini -- enables users to generate high-quality images from textual prompts. However, it has become…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yixin Wan , Arjun Subramonian , Anaelia Ovalle , Zongyu Lin , Ashima Suvarna , Christina Chance , Hritik Bansal , Rebecca Pattichis , Kai-Wei Chang

Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Yabo Zhang , Yuxiang Wei , Xianhui Lin , Zheng Hui , Peiran Ren , Xuansong Xie , Xiangyang Ji , Wangmeng Zuo

Text-to-image generative models excel in creating images from text but struggle with ensuring alignment and consistency between outputs and prompts. This paper introduces TextMatch, a novel framework that leverages multimodal optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Yucong Luo , Mingyue Cheng , Jie Ouyang , Xiaoyu Tao , Qi Liu

Recently, Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated metrics have emerged to evaluate the image-text alignment capabilities of generative models. However, the performance…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Shuhao Han , Haotian Fan , Jiachen Fu , Liang Li , Tao Li , Junhui Cui , Yunqiu Wang , Yang Tai , Jingwei Sun , Chunle Guo , Chongyi Li

Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models optimize for "average" human appeal, they fail to capture the inherent subjectivity of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Anne-Sofie Maerten , Juliane Verwiebe , Shyamgopal Karthik , Ameya Prabhu , Johan Wagemans , Matthias Bethge

While Instruction-based Image Editing (IIE) has achieved significant progress, existing benchmarks pursue task breadth via mixed evaluations. This paradigm obscures a critical failure mode crucial in professional applications: the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yujia Yang , Yuanxiang Wang , Zhenyu Guan , Tiankun Yang , Chenxi Bao , Haopeng Jin , Jinwen Luo , Xinyu Zuo , Lisheng Duan , Haijin Liang , Jin Ma , Xinming Wang , Ruiwen Tao , Hongzhu Yi

While often assumed a gold standard, effective human evaluation of text generation remains an important, open area for research. We revisit this problem with a focus on producing consistent evaluations that are reproducible -- over time and…

Computation and Language · Computer Science 2022-11-02 Daniel Khashabi , Gabriel Stanovsky , Jonathan Bragg , Nicholas Lourie , Jungo Kasai , Yejin Choi , Noah A. Smith , Daniel S. Weld

One challenge in text-to-image (T2I) generation is the inadvertent reflection of culture gaps present in the training data, which signifies the disparity in generated image quality when the cultural elements of the input text are rarely…

Computer Vision and Pattern Recognition · Computer Science 2023-07-07 Bingshuai Liu , Longyue Wang , Chenyang Lyu , Yong Zhang , Jinsong Su , Shuming Shi , Zhaopeng Tu

Text-to-video (T2V) diffusion models have achieved rapid progress, yet their demographic biases, particularly gender bias, remain largely unexplored. We present FairT2V, a training-free debiasing framework for text-to-video generation that…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Haonan Zhong , Wei Song , Tingxu Han , Maurice Pagnucco , Jingling Xue , Yang Song

Text-to-image (T2I) generation aims at producing realistic images corresponding to text descriptions. Generative Adversarial Network (GAN) has proven to be successful in this task. Typical T2I GANs are 2 phase methods that first pretrain an…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Yibin Liu , Jianyu Zhang , Li Zhang , Shijian Li , Gang Pan

Modern generative models have demonstrated the ability to solve challenging mathematical problems. In many real-world settings, however, mathematical solutions must be expressed visually through diagrams, plots, geometric constructions, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ruiyao Liu , Hui Shen , Ping Zhang , Yunta Hsieh , Yifan Zhang , Jing Xu , Sicheng Chen , Junchen Li , Jiawei Lu , Jianing Ma , Jiaqi Mo , Qi Han , Zhen Zhang , Zhongwei Wan , Jing Xiong , Xin Wang , Ziyuan Liu , Hangrui Cao , Ngai Wong

Text-to-Image (T2I) models have achieved remarkable success in generating visual content from text inputs. Although multiple safety alignment strategies have been proposed to prevent harmful outputs, they often lead to overly cautious…

Machine Learning · Computer Science 2025-10-28 Ziheng Cheng , Yixiao Huang , Hui Xu , Somayeh Sojoudi , Xuandong Zhao , Dawn Song , Song Mei

Text-to-video (T2V) generation models have made significant progress in creating visually appealing videos. However, they struggle with generating coherent sequential narratives that require logical progression through multiple events.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Zhengxu Tang , Zizheng Wang , Luning Wang , Zitao Shuai , Chenhao Zhang , Siyu Qian , Yirui Wu , Bohao Wang , Haosong Rao , Zhenyu Yang , Chenwei Wu

Text-to-image (T2I) models exhibit a significant yet under-explored "brand bias", a tendency to generate contents featuring dominant commercial brands from generic prompts, posing ethical and legal risks. We propose CIDER, a novel,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Fangjian Shen , Zifeng Liang , Chao Wang , Wushao Wen

Text-to-image (T2I) generative models have gained increased popularity in the public domain. While boasting impressive user-guided generative abilities, their black-box nature exposes users to intentionally- and intrinsically-biased…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Jordan Vice , Naveed Akhtar , Richard Hartley , Ajmal Mian

Recent advancements in text-to-image (T2I) diffusion models have demonstrated remarkable capabilities in generating high-fidelity images. However, these models often struggle to faithfully render complex user prompts, particularly in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Linqing Wang , Ximing Xing , Yiji Cheng , Zhiyuan Zhao , Donghao Li , Tiankai Hang , Jiale Tao , Qixun Wang , Ruihuang Li , Comi Chen , Xin Li , Mingrui Wu , Xinchi Deng , Shuyang Gu , Chunyu Wang , Qinglin Lu

The rapid development of text-to-image generation has brought rising ethical considerations, especially regarding gender bias. Given a text prompt as input, text-to-image models generate images according to the prompt. Pioneering models…

Computers and Society · Computer Science 2024-08-22 Yankun Wu , Yuta Nakashima , Noa Garcia
‹ Prev 1 3 4 5 6 7 10 Next ›