English
Related papers

Related papers: Text-Driven Image Editing via Learnable Regions

200 papers

Recent large-scale text-guided diffusion models provide powerful image-generation capabilities. Currently, a significant effort is given to enable the modification of these images using text only as means to offer intuitive and versatile…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Linoy Tsaban , Apolinário Passos

Recent advancements in diffusion models have showcased their impressive capacity to generate visually striking images. Nevertheless, ensuring a close match between the generated image and the given prompt remains a persistent challenge. In…

Computer Vision and Pattern Recognition · Computer Science 2023-09-11 Yupeng Zhou , Daquan Zhou , Zuo-Liang Zhu , Yaxing Wang , Qibin Hou , Jiashi Feng

Instruction-based image editing (IIE) aims to modify images according to textual instructions while preserving irrelevant content. Despite recent advances in diffusion transformers, existing methods often suffer from over-editing,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Jingxuan He , Xiyu Wang , Mengyu Zheng , Xiangyu Zeng , Yunke Wang , Chang Xu

Text-to-Image models have introduced a remarkable leap in the evolution of machine learning, demonstrating high-quality synthesis of images from a given text-prompt. However, these powerful pretrained models still lack control handles that…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Andrey Voynov , Kfir Aberman , Daniel Cohen-Or

Text-guided image editing has been allowing users to transform and synthesize images through natural language instructions, offering considerable flexibility. However, most existing image editing models naively attempt to follow all user…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Hyunseung Kim , Chiho Choi , Srikanth Malla , Sai Prahladh Padmanabhan , Saurabh Bagchi , Joon Hee Choi

Users interact with text, image, code, or other editors on a daily basis. However, machine learning models are rarely trained in the settings that reflect the interactivity between users and their editor. This is understandable as training…

Computation and Language · Computer Science 2023-11-14 Felix Faltings , Michel Galley , Baolin Peng , Kianté Brantley , Weixin Cai , Yizhe Zhang , Jianfeng Gao , Bill Dolan

Existing text-to-image diffusion models struggle to synthesize realistic images given dense captions, where each text prompt provides a detailed description for a specific image region. To address this, we propose DenseDiffusion, a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Yunji Kim , Jiyoung Lee , Jin-Hwa Kim , Jung-Woo Ha , Jun-Yan Zhu

This paper addresses the problem of manipulating images using natural language description. Our task aims to semantically modify visual attributes of an object in an image according to the text describing the new visual appearance. Although…

Computer Vision and Pattern Recognition · Computer Science 2018-11-29 Seonghyeon Nam , Yunji Kim , Seon Joo Kim

While language-guided image manipulation has made remarkable progress, the challenge of how to instruct the manipulation process faithfully reflecting human intentions persists. An accurate and comprehensive description of a manipulation…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Yasheng Sun , Yifan Yang , Houwen Peng , Yifei Shen , Yuqing Yang , Han Hu , Lili Qiu , Hideki Koike

We propose a novel algorithm, named Open-Edit, which is the first attempt on open-domain image manipulation with open-vocabulary instructions. It is a challenging task considering the large variation of image domains and the lack of…

Computer Vision and Pattern Recognition · Computer Science 2021-04-22 Xihui Liu , Zhe Lin , Jianming Zhang , Handong Zhao , Quan Tran , Xiaogang Wang , Hongsheng Li

In this paper, we propose a novel controllable text-to-image generative adversarial network (ControlGAN), which can effectively synthesise high-quality images and also control parts of the image generation according to natural language…

Computer Vision and Pattern Recognition · Computer Science 2019-12-20 Bowen Li , Xiaojuan Qi , Thomas Lukasiewicz , Philip H. S. Torr

Text-to-image (T2I) diffusion models, with their impressive generative capabilities, have been adopted for image editing tasks, demonstrating remarkable efficacy. However, due to attention leakage and collision between the cross-attention…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Xingxi Yin , Zhi Li , Jingfeng Zhang , Chenglin Li , Yin Zhang

Text-to-image generation model is able to generate images across a diverse range of subjects and styles based on a single prompt. Recent works have proposed a variety of interaction methods that help users understand the capabilities of…

Human-Computer Interaction · Computer Science 2023-07-19 Seungho Baek , Hyerin Im , Jiseung Ryu , Juhyeong Park , Takyeon Lee

Image editing affords increased control over the aesthetics and content of generated images. Pre-existing works focus predominantly on text-based instructions to achieve desired image modifications, which limit edit precision and accuracy.…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Bowen Li , Yongxin Yang , Steven McDonagh , Shifeng Zhang , Petru-Daniel Tudosiu , Sarah Parisot

Instruction-based image editing enables natural-language control over visual modifications, yet existing models falter under Instruction-Visual Complexity (IV-Complexity), where intricate instructions meet cluttered or ambiguous scenes. We…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Tianyuan Qu , Lei Ke , Xiaohang Zhan , Longxiang Tang , Yuqi Liu , Bohao Peng , Bei Yu , Dong Yu , Jiaya Jia

Free-form text prompts allow users to describe their intentions during image manipulation conveniently. Based on the visual latent space of StyleGAN[21] and text embedding space of CLIP[34], studies focus on how to map these two latent…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Yiming Zhu , Hongyu Liu , Yibing Song , ziyang Yuan , Xintong Han , Chun Yuan , Qifeng Chen , Jue Wang

Text-conditioned image generation models are a prevalent use of AI image synthesis, yet intuitively controlling output guided by an artist remains challenging. Current methods require multiple images and textual prompts for each object to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Shounak Chatterjee

Text-driven localized editing of 3D objects is particularly difficult as locally mixing the original 3D object with the intended new object and style effects without distorting the object's form is not a straightforward process. To address…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Hyeonseop Song , Seokhun Choi , Hoseok Do , Chul Lee , Taehyeong Kim

Text-guided image editing involves modifying a source image based on a language instruction and, typically, requires changes to only small local regions. However, existing approaches generate the entire target image rather than selectively…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Huimin Wu , Xiaojian Ma , Haozhe Zhao , Yanpeng Zhao , Qing Li

Instruction-based image editing has made a great process in using natural human language to manipulate the visual content of images. However, existing models are limited by the quality of the dataset and cannot accurately localize editing…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Tiancheng Li , Jinxiu Liu , Huajun Chen , Qi Liu