English
Related papers

Related papers: ACCORD: Alleviating Concept Coupling through Depen…

200 papers

Text-to-video diffusion transformers encode semantic information unevenly across model depth, which constrains effective concept erasure. We identify a representational bottleneck, termed concept-layer topological alignment, under which…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Yiwei Xie , Ping Liu , Zheng Zhang

Recent advances in generative diffusion models have shown a notable inherent understanding of image style and semantics. In this paper, we leverage the self-attention features from pretrained diffusion networks to transfer the visual…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Yang Zhou , Xu Gao , Zichong Chen , Hui Huang

Image captioning, a fundamental task in vision-language understanding, seeks to generate accurate natural language descriptions for provided images. Current image captioning approaches heavily rely on high-quality image-caption pairs, which…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Chuanyang Jin

Learning from human feedback has been shown to improve text-to-image models. These techniques first learn a reward function that captures what humans care about in the task and then improve the models based on the learned reward function.…

Inverse problems appear in many applications, such as image deblurring and inpainting. The common approach to address them is to design a specific algorithm for each problem. The Plug-and-Play (P&P) framework, which has been recently…

Computer Vision and Pattern Recognition · Computer Science 2018-10-12 Tom Tirer , Raja Giryes

Despite recent advancements, the field of text-to-image synthesis still suffers from lack of fine-grained control. Using only text, it remains challenging to deal with issues such as concept coherence and concept contamination. We propose a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Rodrigo Mello , Filipe Calegario , Geber Ramalho

We propose a way of learning disentangled content-style representation of image, allowing us to extrapolate images to any style as well as interpolate between any pair of styles. By augmenting data set in a supervised setting and imposing…

Computer Vision and Pattern Recognition · Computer Science 2021-12-01 Sailun Xu , Jiazhi Zhang , Jiamei Liu

Art reinterpretation is the practice of creating a variation of a reference work, making a paired artwork that exhibits a distinct artistic style. We ask if such an image pair can be used to customize a generative model to capture the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Maxwell Jones , Sheng-Yu Wang , Nupur Kumari , David Bau , Jun-Yan Zhu

Cell counting in microscopy images is vital in medicine and biology but extremely tedious and time-consuming to perform manually. While automated methods have advanced in recent years, state-of-the-art approaches tend to increasingly…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Zixuan Zheng , Yilei Shi , Chunlei Li , Jingliang Hu , Xiao Xiang Zhu , Lichao Mou

Text-to-image diffusion models have demonstrated remarkable capabilities in creating images highly aligned with user prompts, yet their proclivity for memorizing training set images has sparked concerns about the originality of the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Chen Chen , Daochang Liu , Mubarak Shah , Chang Xu

Object detection plays an important role in current solutions to vision and language tasks like image captioning and visual question answering. However, popular models like Faster R-CNN rely on a costly process of annotating ground-truths…

Computation and Language · Computer Science 2019-09-06 Soravit Changpinyo , Bo Pang , Piyush Sharma , Radu Soricut

Dense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-09 Linjie Yang , Kevin Tang , Jianchao Yang , Li-Jia Li

Temporal grounding aims to localize temporal boundaries within untrimmed videos by language queries, but it faces the challenge of two types of inevitable human uncertainties: query uncertainty and label uncertainty. The two uncertainties…

Computer Vision and Pattern Recognition · Computer Science 2021-06-25 Hao Zhou , Chongyang Zhang , Yan Luo , Yanjun Chen , Chuanping Hu

Recently, the multimedia community has witnessed the rise of diffusion models trained on large-scale multi-modal data for visual content creation, particularly in the field of text-to-image generation. In this paper, we propose a new task…

Computer Vision and Pattern Recognition · Computer Science 2023-11-10 Jingwen Chen , Yingwei Pan , Ting Yao , Tao Mei

Customized generation aims to incorporate a novel concept into a pre-trained text-to-image model, enabling new generations of the concept in novel contexts guided by textual prompts. However, customized generation suffers from an inherent…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Jian Jin , Yang Shen , Zhenyong Fu , Jian Yang

Synthesizing visually impressive images that seamlessly align both text prompts and specific artistic styles remains a significant challenge in Text-to-Image (T2I) diffusion models. This paper introduces StyleBlend, a method designed to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Zichong Chen , Shijin Wang , Yang Zhou

We observe that the mapping between an image's representation in one model to its representation in another can be learned surprisingly well with just a linear layer, even across diverse models. Building on this observation, we propose…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Mazda Moayeri , Keivan Rezaei , Maziar Sanjabi , Soheil Feizi

Personalized image generation requires text-to-image generative models that capture the core features of a reference subject to allow for controlled generation across different contexts. Existing methods face challenges due to complex…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Emanuele Aiello , Umberto Michieli , Diego Valsesia , Mete Ozay , Enrico Magli

Diffusion-based image editing is a composite process of preserving the source image content and generating new content or applying modifications. While current editing approaches have made improvements under text guidance, most of them have…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Tianrui Huang , Pu Cao , Lu Yang , Chun Liu , Mengjie Hu , Zhiwei Liu , Qing Song

Recent text-to-image diffusion models are able to learn and synthesize images containing novel, personalized concepts (e.g., their own pets or specific items) with just a few examples for training. This paper tackles two interconnected…

Computer Vision and Pattern Recognition · Computer Science 2024-02-26 Chun-Hsiao Yeh , Ta-Ying Cheng , He-Yen Hsieh , Chuan-En Lin , Yi Ma , Andrew Markham , Niki Trigoni , H. T. Kung , Yubei Chen