English
Related papers

Related papers: ConceptPrism: Concept Disentanglement in Personali…

200 papers

Erasing concepts from large-scale text-to-image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infringement, offensive content, and privacy violations. In scalable applications,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ouxiang Li , Yuan Wang , Xinting Hu , Houcheng Jiang , Yanbin Hao , Fuli Feng

Diffusion customization methods have achieved impressive results with only a minimal number of user-provided images. However, existing approaches customize concepts collectively, whereas real-world applications often require sequential…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Zirun Guo , Tao Jin

Text-to-Image diffusion models can produce undesirable content that necessitates concept erasure. However, existing methods struggle with under-erasure, leaving residual traces of targeted concepts, or over-erasure, mistakenly eliminating…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Yuyang Xue , Edward Moroshko , Feng Chen , Jingyu Sun , Steven McDonagh , Sotirios A. Tsaftaris

Text-to-image (T2I) diffusion models often inadvertently generate unwanted concepts such as watermarks and unsafe images. These concepts, termed as the "implicit concepts", could be unintentionally learned during training and then be…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Zhili Liu , Kai Chen , Yifan Zhang , Jianhua Han , Lanqing Hong , Hang Xu , Zhenguo Li , Dit-Yan Yeung , James Kwok

Concept erasing has recently emerged as an effective paradigm to prevent text-to-image diffusion models from generating visually undesirable or even harmful content. However, current removal methods heavily rely on manually crafted text…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Feiran Li , Qianqian Xu , Shilong Bao , Zhiyong Yang , Xiaochun Cao , Qingming Huang

Removing undesired concepts from large-scale text-to-image (T2I) and text-to-video (T2V) diffusion models while preserving overall generative quality remains a major challenge, particularly as modern models such as Stable Diffusion v3,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhaoxin Fan , Nanxiang Jiang , Daiheng Gao , Shiji Zhou , Wenjun Wu

This paper explores the possibility of learning custom tokens for representing new concepts in Vision-Language Models (VLMs). Our aim is to learn tokens that can be effective for both discriminative and generative tasks while composing well…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Pramuditha Perera , Matthew Trager , Luca Zancato , Alessandro Achille , Stefano Soatto

Personalized text-to-image generation aims to create images tailored to user-defined concepts and textual descriptions. Balancing the fidelity of the learned concept with its ability for generation in various contexts presents a significant…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Vera Soboleva , Maksim Nakhodnov , Aibek Alanov

Text-guided diffusion models have revolutionized generative tasks by producing high-fidelity content from text descriptions. They have also enabled an editing paradigm where concepts can be replaced through text conditioning (e.g., a dog to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Chao Huang , Susan Liang , Yunlong Tang , Yapeng Tian , Anurag Kumar , Chenliang Xu

Studies have been conducted to prevent specific concepts from being generated from pretrained text-to-image generative models, achieving concept erasure in various ways. However, the performance evaluation of these studies is still largely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Masane Fuchi , Tomohiro Takagi

In this paper, we investigate the problem of learning disentangled representations. Given a pair of images sharing some attributes, we aim to create a low-dimensional representation which is split into two parts: a shared representation…

Machine Learning · Statistics 2019-12-10 Eduardo Hugo Sanchez , Mathieu Serrurier , Mathias Ortner

We present TokenVerse -- a method for multi-concept personalization, leveraging a pre-trained text-to-image diffusion model. Our framework can disentangle complex visual elements and attributes from as little as a single image, while…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Daniel Garibi , Shahar Yadin , Roni Paiss , Omer Tov , Shiran Zada , Ariel Ephrat , Tomer Michaeli , Inbar Mosseri , Tali Dekel

Text-to-image customization, which aims to synthesize text-driven images for the given subjects, has recently revolutionized content creation. Existing works follow the pseudo-word paradigm, i.e., represent the given subjects as…

Computer Vision and Pattern Recognition · Computer Science 2024-03-04 Mengqi Huang , Zhendong Mao , Mingcong Liu , Qian He , Yongdong Zhang

Text-to-image diffusion models have achieved remarkable progress in generating diverse and realistic images from textual descriptions. However, they still struggle with personalization, which requires adapting a pretrained model to depict…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Seoyun Yang , Gihoon Kim , Taesup Kim

Distributed representations provide a vector space that captures meaningful relationships between data instances. The distributed nature of these representations, however, entangles together multiple attributes or concepts of data instances…

Machine Learning · Computer Science 2023-12-04 Somnath Basu Roy Chowdhury , Nicholas Monath , Avinava Dubey , Amr Ahmed , Snigdha Chaturvedi

Concept erasure techniques have been widely deployed in T2I diffusion models to prevent inappropriate content generation for safety and copyright considerations. However, as models evolve to next-generation architectures like Flux,…

Machine Learning · Computer Science 2025-10-07 Daiheng Gao , Nanxiang Jiang , Andi Zhang , Shilin Lu , Yufei Tang , Wenbo Zhou , Weiming Zhang , Zhaoxin Fan

Compositionality is a critical capability in Text-to-Image (T2I) models, as it reflects their ability to understand and combine multiple concepts from text descriptions. Existing evaluations of compositional capability rely heavily on…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Xindi Wu , Dingli Yu , Yangsibo Huang , Olga Russakovsky , Sanjeev Arora

Image-to-image translation (i2i) networks suffer from entanglement effects in presence of physics-related phenomena in target domain (such as occlusions, fog, etc), lowering altogether the translation quality, controllability and…

Computer Vision and Pattern Recognition · Computer Science 2023-04-28 Fabio Pizzati , Pietro Cerri , Raoul de Charette

Text-to-image person re-identification (TIReID) aims to retrieve person images from a large gallery given free-form textual descriptions. TIReID is challenging due to the substantial modality gap between visual appearances and textual…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Giyeol Kim , Chanho Eom

Vision-language co-embedding networks, such as CLIP, provide a latent embedding space with semantic information that is useful for downstream tasks. We hypothesize that the embedding space can be disentangled to separate the information on…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Zhi Li , Hau Phan , Matthew Emigh , Austin J. Brockmeier