中文
相关论文

相关论文: Order Is Not Layout: Order-to-Space Bias in Image …

200 篇论文

With the growing adoption of Text-to-Image (TTI) systems, the social biases of these models have come under increased scrutiny. Herein we conduct a systematic investigation of one such source of bias for diffusion models: embedding spaces.…

机器学习 · 计算机科学 2024-09-17 Sahil Kuchlous , Marvin Li , Jeffrey G. Wang

A wide range of applications require learning image generation models whose latent space effectively captures the high-level factors of variation present in the data distribution. The extent to which a model represents such variations…

计算机视觉与模式识别 · 计算机科学 2021-11-09 Avinandan Bose , Aniket Das , Yatin Dandi , Piyush Rai

Despite recent impressive results on single-object and single-domain image generation, the generation of complex scenes with multiple objects remains challenging. In this paper, we start with the idea that a model must be able to understand…

计算机视觉与模式识别 · 计算机科学 2020-12-04 Tristan Sylvain , Pengchuan Zhang , Yoshua Bengio , R Devon Hjelm , Shikhar Sharma

It is common in graphic design humans visually arrange various elements according to their design intent and semantics. For example, a title text almost always appears on top of other elements in a document. In this work, we generate…

计算机视觉与模式识别 · 计算机科学 2021-08-03 Kotaro Kikuchi , Edgar Simo-Serra , Mayu Otani , Kota Yamaguchi

Recent approaches have achieved great success in image generation from structured inputs, e.g., semantic segmentation, scene graph or layout. Although these methods allow specification of objects and their locations at image-level, they…

计算机视觉与模式识别 · 计算机科学 2020-08-28 Ke Ma , Bo Zhao , Leonid Sigal

We propose a novel training-free image generation algorithm that precisely controls the occlusion relationships between objects in an image. Existing image generation methods typically rely on prompts to influence occlusion, which often…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Xiaohang Zhan , Dingming Liu

Positional bias - where models overemphasize certain positions regardless of content - has been shown to negatively impact model performance across various tasks. While recent research has extensively examined positional bias in text…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Kebin Wu , Fatima Albreiki

As text-to-image systems continue to grow in popularity with the general public, questions have arisen about bias and diversity in the generated images. Here, we investigate properties of images generated in response to prompts which are…

计算机与社会 · 计算机科学 2023-02-15 Kathleen C. Fraser , Svetlana Kiritchenko , Isar Nejadgholi

Vision-Language Models have demonstrated remarkable capabilities in understanding visual content, yet systematic biases in their spatial processing remain largely unexplored. This work identifies and characterizes a systematic spatial…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Aryan Chaudhary , Sanchit Goyal , Pratik Narang , Dhruv Kumar

The rapid development of text-to-image generation has brought rising ethical considerations, especially regarding gender bias. Given a text prompt as input, text-to-image models generate images according to the prompt. Pioneering models…

计算机与社会 · 计算机科学 2024-08-22 Yankun Wu , Yuta Nakashima , Noa Garcia

Controllable text-to-image generation synthesizes visual text and objects in images with certain conditions, which are frequently applied to emoji and poster generation. Visual text rendering and layout-to-image generation tasks have been…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Xiaoran Zhao , Tianhao Wu , Yu Lai , Zhiliang Tian , Zhen Huang , Yahui Liu , Zejiang He , Dongsheng Li

Matching images and sentences demands a fine understanding of both modalities. In this paper, we propose a new system to discriminatively embed the image and text to a shared visual-textual space. In this field, most existing works apply…

计算机视觉与模式识别 · 计算机科学 2021-07-28 Zhedong Zheng , Liang Zheng , Michael Garrett , Yi Yang , Mingliang Xu , Yi-Dong Shen

Fine-grained open-set recognition (FineOSR) aims to recognize images belonging to classes with subtle appearance differences while rejecting images of unknown classes. A recent trend in OSR shows the benefit of generative models to…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Wentao Bao , Qi Yu , Yu Kong

Recent progress in Text-to-Image (T2I) generative models has enabled high-quality image generation. As performance and accessibility increase, these models are gaining significant attraction and popularity: ensuring their fairness and…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Moreno D'Incà , Elia Peruzzo , Massimiliano Mancini , Xingqian Xu , Humphrey Shi , Nicu Sebe

A layout to image (L2I) generation model aims to generate a complicated image containing multiple objects (things) against natural background (stuff), conditioned on a given layout. Built upon the recent advances in generative adversarial…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Sen He , Wentong Liao , Michael Ying Yang , Yongxin Yang , Yi-Zhe Song , Bodo Rosenhahn , Tao Xiang

Recent training-free layout-to-image diffusion models have demonstrated remarkable performance in generating high-quality images with controllable layouts. These models follow a one-stage framework: Encouraging the model to focus the…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Linhao Huang , Jing Yu

Bias amplification is a phenomenon in which models exacerbate biases or stereotypes present in the training data. In this paper, we study bias amplification in the text-to-image domain using Stable Diffusion by comparing gender ratios in…

机器学习 · 计算机科学 2023-11-16 Preethi Seshadri , Sameer Singh , Yanai Elazar

Large text-to-image diffusion models have impressive capabilities in generating photorealistic images from text prompts. How to effectively guide or control these powerful models to perform different downstream tasks becomes an important…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Zeju Qiu , Weiyang Liu , Haiwen Feng , Yuxuan Xue , Yao Feng , Zhen Liu , Dan Zhang , Adrian Weller , Bernhard Schölkopf

Existing open-set recognition (OSR) studies typically assume that each image contains only one class label, with the unknown test set (negative) having a disjoint label space from the known test set (positive), a scenario referred to as…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Xu Yin , Fei Pan , Guoyuan An , Yuchi Huo , Zixuan Xie , Sung-Eui Yoon

In text-to-image generation tasks, the advancements of diffusion models have facilitated the fidelity of generated results. However, these models encounter challenges when processing text prompts containing multiple entities and attributes.…

计算与语言 · 计算机科学 2024-04-23 Yihang Wu , Xiao Cao , Kaixin Li , Zitan Chen , Haonan Wang , Lei Meng , Zhiyong Huang