English
Related papers

Related papers: Identifying Prompted Artist Names from Generated I…

200 papers

While text-to-image synthesis currently enjoys great popularity among researchers and the general public, the security of these models has been neglected so far. Many text-guided image generation models rely on pre-trained text encoders…

Machine Learning · Computer Science 2023-08-10 Lukas Struppek , Dominik Hintersdorf , Kristian Kersting

With the increasing use of image generation technology, understanding its social biases, including gender bias, is essential. This paper presents a large-scale study on gender bias in text-to-image (T2I) models, focusing on everyday…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Leander Girrbach , Stephan Alaniz , Genevieve Smith , Zeynep Akata

With the spread of the use of Text2Img diffusion models such as DALL-E 2, Imagen, Mid Journey and Stable Diffusion, one challenge that artists face is selecting the right prompts to achieve the desired artistic output. We present techniques…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Sam Witteveen , Martin Andrews

Despite the remarkable progress of image captioning, existing captioners typically lack the controllable capability to generate desired image captions, e.g., describing the image in a rough or detailed manner, in a factual or emotional…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Ning Wang , Jiahao Xie , Jihao Wu , Mingbo Jia , Linlin Li

Text-to-image generation models have grown in popularity due to their ability to produce high-quality images from a text prompt. One use for this technology is to enable the creation of more accessible art creation software. In this paper,…

Human-Computer Interaction · Computer Science 2023-09-06 Atieh Taheri , Mohammad Izadi , Gururaj Shriram , Negar Rostamzadeh , Shaun Kane

Text-to-image diffusion models have achieved remarkable performance in image synthesis, while the text interface does not always provide fine-grained control over certain image factors. For instance, changing a single token in the text can…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Chen Wu , Fernando De la Torre

Training data memorization in language models impacts model capability (generalization) and safety (privacy risk). This paper focuses on analyzing prompts' impact on detecting the memorization of 6 masked language model-based named entity…

Computation and Language · Computer Science 2024-05-07 Yuxi Xia , Anastasiia Sedova , Pedro Henrique Luz de Araujo , Vasiliki Kougia , Lisa Nußbaumer , Benjamin Roth

Current text-to-image (T2I) generation models achieve promising results, but they fail on the scenarios where the knowledge implied in the text prompt is uncertain. For example, a T2I model released in February would struggle to generate a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Chuanhao Li , Jianwen Sun , Yukang Feng , Mingliang Zhai , Yifan Chang , Kaipeng Zhang

Conditional generative models such as DALL-E and Stable Diffusion generate images based on a user-defined text, the prompt. Finding and refining prompts that produce a desired image has become the art of prompt engineering. Generative…

Information Retrieval · Computer Science 2023-01-24 Niklas Deckers , Maik Fröbe , Johannes Kiesel , Gianluca Pandolfo , Christopher Schröder , Benno Stein , Martin Potthast

Synthesis of digital artifacts conditioned on user prompts has become an important paradigm facilitating an explosion of use cases with generative AI. However, such models often fail to connect the generated outputs and desired target…

Machine Learning · Computer Science 2026-04-15 Melvin Wong , Yew-Soon Ong , Abhishek Gupta , Kavitesh K. Bali , Caishun Chen

Text-to-image models are enabling efficient design space exploration, rapidly generating images from text prompts. However, many generative AI tools are imperfect for product design applications as they are not built for the goals and…

Human-Computer Interaction · Computer Science 2025-01-22 Leah Chong , I-Ping Lo , Jude Rayan , Steven Dow , Faez Ahmed , Ioanna Lykourentzou

Despite the impressive advances in text-to-image models, they often struggle to effectively compose complex scenes with multiple objects, displaying various attributes and relationships. To address this challenge, we present…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Kaiyi Huang , Chengqi Duan , Kaiyue Sun , Enze Xie , Zhenguo Li , Xihui Liu

Facial images have extensive practical applications. Although the current large-scale text-image diffusion models exhibit strong generation capabilities, it is challenging to generate the desired facial images using only text prompt. Image…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Dawei Dai , Mingming Jia , Yinxiu Zhou , Hang Xing , Chenghang Li

Recent text-to-image (T2I) generation models have advanced significantly, enabling the creation of high-fidelity images from textual prompts. However, existing evaluation benchmarks primarily focus on the explicit alignment between…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Wenchao Zhang , Jiahe Tian , Runze He , Jizhong Han , Jiao Dai , Miaomiao Feng , Wei Mi , Xiaodan Zhang

Text-to-audio (TTA) systems are rapidly transforming music creation and distribution, with platforms like Udio and Suno generating thousands of tracks daily and integrating into mainstream music platforms and ecosystems. These systems,…

Sound · Computer Science 2025-11-24 Guilherme Coelho

Text-to-image (T2I) systems increasingly rely on upstream prompters, either humans or multimodal large language models (MLLMs), to translate user intent into detailed prompts. Yet current benchmarks fix the prompt and only evaluate T2I…

Artificial Intelligence · Computer Science 2026-05-22 Hanjun Luo , Zhimu Huang , Sylvia Chung , Yiran Wang , Yingbin Jin , Jialin Li , Jiang Li , Xinfeng Li , Hanan Salam

Text-to-image models take a sentence (i.e., prompt) and generate images associated with this input prompt. These models have created award wining-art, videos, and even synthetic datasets. However, text-to-image (T2I) models can generate…

Computation and Language · Computer Science 2023-06-12 Alexander Lin , Lucas Monteiro Paes , Sree Harsha Tanneru , Suraj Srinivas , Himabindu Lakkaraju

Recent advances in text-to-image (T2I) generation have achieved remarkable visual outcomes through large-scale rectified flow models. However, how these models behave under long prompts remains underexplored. Long prompts encode rich…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Bo-Kai Ruan , Teng-Fang Hsiao , Ling Lo , Yi-Lun Wu , Hong-Han Shuai

While text-to-image (T2I) models can synthesize high-quality images, their performance degrades significantly when prompted with novel or out-of-distribution (OOD) entities due to inherent knowledge cutoffs. We introduce World-To-Image, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Moo Hyun Son , Jintaek Oh , Sun Bin Mun , Jaechul Roh , Sehyun Choi

While an image is worth more than a thousand words, only a few provide crucial information for a given task and thus should be focused on. In light of this, ideal text-to-image (T2I) retrievers should prioritize specific visual attributes…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Siting Li , Xiang Gao , Simon Shaolei Du
‹ Prev 1 3 4 5 6 7 10 Next ›