中文
相关论文

相关论文: Magnet: We Never Know How Text-to-Image Diffusion …

200 篇论文

Text-to-image diffusion models rely on massive, web-scale datasets. Training them from scratch is computationally expensive, and as a result, developers often prefer to make incremental updates to existing models. These updates often…

机器学习 · 计算机科学 2025-09-29 Vinith M. Suriyakumar , Rohan Alur , Ayush Sekhari , Manish Raghavan , Ashia C. Wilson

There has been a significant progress in text conditional image generation models. Recent advancements in this field depend not only on improvements in model structures, but also vast quantities of text-image paired datasets. However,…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Seungdae Han , Joohee Kim

Recent advances in text-to-image (T2I) diffusion models have significantly improved the quality of generated images. However, providing efficient control over individual subjects, particularly the attributes characterizing them, remains a…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Stefan Andreas Baumann , Felix Krause , Michael Neumayr , Nick Stracke , Melvin Sevi , Vincent Tao Hu , Björn Ommer

While vision-language models (VLMs) have achieved remarkable performance improvements recently, there is growing evidence that these models also posses harmful biases with respect to social attributes such as gender and race. Prior studies…

计算机视觉与模式识别 · 计算机科学 2023-10-05 Phillip Howard , Avinash Madasu , Tiep Le , Gustavo Lujan Moreno , Vasudev Lal

Recent breakthroughs of transformer-based diffusion models, particularly with Multimodal Diffusion Transformers (MMDiT) driven models like FLUX and Qwen Image, have facilitated thrilling experiences in text-to-image generation and editing.…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Binglei Li , Mengping Yang , Zhiyu Tan , Junping Zhang , Hao Li

We learn visual features by captioning images with an image-conditioned masked diffusion language model, a formulation we call masked diffusion captioning (MDC). During training, text tokens in each image-caption pair are masked at a…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Chao Feng , Zihao Wei , Andrew Owens

Text-to-image synthesis has recently attracted widespread attention due to rapidly improving quality and numerous practical applications. However, the language understanding capabilities of text-to-image models are still poorly understood,…

计算机视觉与模式识别 · 计算机科学 2023-10-16 Anton Baryshnikov , Max Ryabinin

The diffusion model has demonstrated superior performance in synthesizing diverse and high-quality images for text-guided image translation. However, there remains room for improvement in both the formulation of text prompts and the…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Qi Si , Bo Wang , Zhao Zhang

Recent advancements in diffusion-based text-to-image (T2I) models have enabled the generation of high-quality and photorealistic images from text. However, they often exhibit societal biases related to gender, race, and socioeconomic…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Jeonghoon Park , Juyoung Lee , Chaeyeon Chung , Jaeseong Lee , Jaegul Choo , Jindong Gu

Text-to-image generative models often exhibit bias related to sensitive attributes. However, current research tends to focus narrowly on single-object prompts with limited contextual diversity. In reality, each object or attribute within a…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Jeng-Lin Li , Ming-Ching Chang , Wei-Chao Chen

Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without retraining. Existing methods either require computationally…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Agata Żywot , Iason Skylitsis , Thijmen Nijdam , Zoe Tzifa-Kratira , Derck Prinzhorn , Konrad Szewczyk , Aritra Bhowmik

As large-scale text-to-image generation models have made remarkable progress in the field of text-to-image generation, many fine-tuning methods have been proposed. However, these models often struggle with novel objects, especially with…

计算机视觉与模式识别 · 计算机科学 2024-01-30 Jianxiang Lu , Cong Xie , Hui Guo

Referring image segmentation is a challenging task that involves generating pixel-wise segmentation masks based on natural language descriptions. The complexity of this task increases with the intricacy of the sentences provided. Existing…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Hai Nguyen-Truong , E-Ro Nguyen , Tuan-Anh Vu , Minh-Triet Tran , Binh-Son Hua , Sai-Kit Yeung

Text encoders in diffusion models have rapidly evolved, transitioning from CLIP to T5-XXL. Although this evolution has significantly enhanced the models' ability to understand complex prompts and generate text, it also leads to a…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Lifu Wang , Daqing Liu , Xinchen Liu , Xiaodong He

Blending visual and textual concepts into a new visual concept is a unique and powerful trait of human beings that can fuel creativity. However, in practice, cross-modal conceptual blending for humans is prone to cognitive biases, like…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Wonwoong Cho , Yanxia Zhang , Yan-Ying Chen , David I. Inouye

In the last two years, text-to-image diffusion models have become extremely popular. As their quality and usage increase, a major concern has been the need for better output control. In addition to prompt engineering, one effective method…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Clément Bonnet , Ariel N. Lee , Franck Wertel , Antoine Tamano , Tanguy Cizain , Pablo Ducru

Customized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suffer from overfitting issues and entangle the…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Yufei Cai , Yuxiang Wei , Zhilong Ji , Jinfeng Bai , Hu Han , Wangmeng Zuo

Due to the high potential for abuse of GenAI systems, the task of detecting synthetic images has recently become of great interest to the research community. Unfortunately, existing image-space detectors quickly become obsolete as new…

计算机视觉与模式识别 · 计算机科学 2024-06-14 George Cazenavette , Avneesh Sud , Thomas Leung , Ben Usman

Recent advancements in text-to-image generation models have dramatically enhanced the generation of photorealistic images from textual prompts, leading to an increased interest in personalized text-to-image applications, particularly in…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Xierui Wang , Siming Fu , Qihan Huang , Wanggui He , Hao Jiang

Models for text-to-image synthesis, such as DALL-E~2 and Stable Diffusion, have recently drawn a lot of interest from academia and the general public. These models are capable of producing high-quality images that depict a variety of…

计算机视觉与模式识别 · 计算机科学 2024-01-10 Lukas Struppek , Dominik Hintersdorf , Felix Friedrich , Manuel Brack , Patrick Schramowski , Kristian Kersting