English
Related papers

Related papers: Stylebreeder: Exploring and Democratizing Artistic…

200 papers

A common and controversial use of text-to-image models is to generate pictures by explicitly naming artists, such as "in the style of Greg Rutkowski". We introduce a benchmark for prompted-artist recognition: predicting which artist names…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Grace Su , Sheng-Yu Wang , Aaron Hertzmann , Eli Shechtman , Jun-Yan Zhu , Richard Zhang

State-of-the-art visual generative AI tools hold immense potential to assist users in the early ideation stages of creative tasks -- offering the ability to generate (rather than search for) novel and unprecedented (instead of existing)…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Evans Xu Han , Alice Qian Zhang , Haiyi Zhu , Hong Shen , Paul Pu Liang , Jane Hsieh

Image editing aims to edit the given synthetic or real image to meet the specific requirements from users. It is widely studied in recent years as a promising and challenging field of Artificial Intelligence Generative Content (AIGC).…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Xincheng Shuai , Henghui Ding , Xingjun Ma , Rongcheng Tu , Yu-Gang Jiang , Dacheng Tao

Taking advantage of the many recent advances in deep learning, text-to-image generative models currently have the merit of attracting the general public attention. Two of these models, DALL-E 2 and Imagen, have demonstrated that highly…

Computer Vision and Pattern Recognition · Computer Science 2022-09-23 Robin Zbinden

We present StyleBabel, a unique open access dataset of natural language captions and free-form tags describing the artistic style of over 135K digital artworks, collected via a novel participatory method from experts studying at specialist…

Computer Vision and Pattern Recognition · Computer Science 2022-03-14 Dan Ruta , Andrew Gilbert , Pranav Aggarwal , Naveen Marri , Ajinkya Kale , Jo Briggs , Chris Speed , Hailin Jin , Baldo Faieta , Alex Filipkowski , Zhe Lin , John Collomosse

As generative AI becomes more prevalent, it is important to study how human users interact with such models. In this work, we investigate how people use text-to-image models to generate desired target images. To study this interaction, we…

Artificial Intelligence · Computer Science 2024-06-18 Kailas Vodrahalli , James Zou

Generative text-to-image models are typically trained on large-scale web-scraped datasets that include diverse visual content such as copyrighted and stylistically distinctive artworks, raising concerns about ownership, attribution, and the…

Machine Learning · Computer Science 2026-05-19 Ninad Joshi , Ashutosh Ranjan , Vivek Srivastava , Shirish Karande

Storyboarding is an established method for designing user experiences. Generative AI can support this process by helping designers quickly create visual narratives. However, existing tools only focus on accurate text-to-image generation.…

Human-Computer Interaction · Computer Science 2024-07-11 Zhaohui Liang , Xiaoyu Zhang , Kevin Ma , Zhao Liu , Xipei Ren , Kosa Goucher-Lambert , Can Liu

In this study, we aim to enhance the capabilities of diffusion-based text-to-image (T2I) generation models by integrating diverse modalities beyond textual descriptions within a unified framework. To this end, we categorize widely used…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Sungnyun Kim , Junsoo Lee , Kibeom Hong , Daesik Kim , Namhyuk Ahn

Text-guided synthesis of images has made a giant leap towards becoming a mainstream phenomenon. With text-to-image generation systems, anybody can create digital images and artworks. This provokes the question of whether text-to-image…

Human-Computer Interaction · Computer Science 2022-11-01 Jonas Oppenlaender

We investigate bias trends in text-to-image generative models over time, focusing on the increasing availability of models through open platforms like Hugging Face. While these platforms democratize AI, they also facilitate the spread of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Jordan Vice , Naveed Akhtar , Richard Hartley , Ajmal Mian

Text-to-image generation model is able to generate images across a diverse range of subjects and styles based on a single prompt. Recent works have proposed a variety of interaction methods that help users understand the capabilities of…

Human-Computer Interaction · Computer Science 2023-07-19 Seungho Baek , Hyerin Im , Jiseung Ryu , Juhyeong Park , Takyeon Lee

Text-to-image generation models have seen considerable advancement, catering to the increasing interest in personalized image creation. Current customization techniques often necessitate users to provide multiple images (typically 3-5) for…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Linhao Zhong , Yan Hong , Wentao Chen , Binglin Zhou , Yiyi Zhang , Jianfu Zhang , Liqing Zhang

Text-to-image diffusion models have emerged as powerful tools for high-quality image generation and editing. Many existing approaches rely on text prompts as editing guidance. However, these methods are constrained by the need for manual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Yuanyuan Chang , Yinghua Yao , Tao Qin , Mengmeng Wang , Ivor Tsang , Guang Dai

Large-scale Text-to-image Generation Models (LTGMs) (e.g., DALL-E), self-supervised deep learning models trained on a huge dataset, have demonstrated the capacity for generating high-quality open-domain images from multi-modal input.…

Human-Computer Interaction · Computer Science 2023-02-17 Hyung-Kwon Ko , Gwanmo Park , Hyeon Jeon , Jaemin Jo , Juho Kim , Jinwook Seo

Text-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of…

Human-Computer Interaction · Computer Science 2023-09-29 Vivian Liu , Lydia B. Chilton

Copyright law confers upon creators the exclusive rights to reproduce, distribute, and monetize their creative works. However, recent progress in text-to-image generation has introduced formidable challenges to copyright enforcement. These…

Computer Vision and Pattern Recognition · Computer Science 2024-06-24 Rui Ma , Qiang Zhou , Yizhu Jin , Daquan Zhou , Bangjun Xiao , Xiuyu Li , Yi Qu , Aishani Singh , Kurt Keutzer , Jingtong Hu , Xiaodong Xie , Zhen Dong , Shanghang Zhang , Shiji Zhou

Artistic painting has achieved significant progress during recent years. Using an autoencoder to connect the original images with compressed latent spaces and a cross attention enhanced U-Net as the backbone of diffusion, latent diffusion…

Computer Vision and Pattern Recognition · Computer Science 2022-10-03 Xianchao Wu

Text-to-image (TTI) systems, particularly those utilizing open-source frameworks, have become increasingly prevalent in the production of Artificial Intelligence (AI)-generated visuals. While existing literature has explored various…

Human-Computer Interaction · Computer Science 2024-08-29 Maria-Teresa De Rosa Palmini , Laura Wagner , Eva Cetinic

Recent advancements in text-to-3D generation have significantly contributed to the automation and democratization of 3D content creation. Building upon these developments, we aim to address the limitations of current methods in blending…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Yeongtak Oh , Jooyoung Choi , Yongsung Kim , Minjun Park , Chaehun Shin , Sungroh Yoon