English
Related papers

Related papers: MagicBrush: A Manually Annotated Dataset for Instr…

200 papers

Enhancing AI systems to perform tasks following human instructions can significantly boost productivity. In this paper, we present InstructP2P, an end-to-end framework for 3D shape editing on point clouds, guided by high-level textual…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Jiale Xu , Xintao Wang , Yan-Pei Cao , Weihao Cheng , Ying Shan , Shenghua Gao

Existing text-guided image manipulation methods aim to modify the appearance of the image or to edit a few objects in a virtual or simple scenario, which is far from practical application. In this work, we study a novel task on text-guided…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Jianan Wang , Guansong Lu , Hang Xu , Zhenguo Li , Chunjing Xu , Yanwei Fu

Text-driven image editing has achieved remarkable success in following single instructions. However, real-world scenarios often involve complex, multi-step instructions, particularly ``chain'' instructions where operations are…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Chenglin Wang , Yucheng Zhou , Qianning Wang , Zhe Wang , Kai Zhang

Retouching can significantly elevate the visual appeal of photos, but many casual photographers lack the expertise to do this well. To address this problem, previous works have proposed automatic retouching systems based on supervised…

Graphics · Computer Science 2018-02-09 Yuanming Hu , Hao He , Chenxi Xu , Baoyuan Wang , Stephen Lin

Due to the availability of increasingly large amounts of visual data, there is a growing need for tools that can help users find relevant images. While existing tools can perform image retrieval based on similarity or metadata, they fall…

Human-Computer Interaction · Computer Science 2024-01-22 Celeste Barnaby , Qiaochu Chen , Chenglong Wang , Isil Dillig

Synthesizing high-resolution images from intricate, domain-specific information remains a significant challenge in generative modeling, particularly for applications in large-image domains such as digital histopathology and remote sensing.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Minh-Quan Le , Alexandros Graikos , Srikar Yellapragada , Rajarsi Gupta , Joel Saltz , Dimitris Samaras

Recent diffusion-based image editing methods have significantly advanced text-guided tasks but often struggle to interpret complex, indirect instructions. Moreover, current models frequently suffer from poor identity preservation,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Chun-Hsiao Yeh , Yilin Wang , Nanxuan Zhao , Richard Zhang , Yuheng Li , Yi Ma , Krishna Kumar Singh

We introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions. Our approach empowers video manipulation guided by natural language directives, eliminating the need for…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Bosheng Qin , Juncheng Li , Siliang Tang , Tat-Seng Chua , Yueting Zhuang

There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Jiafeng Liang , Shixin Jiang , Zekun Wang , Haojie Pan , Zerui Chen , Zheng Chu , Ming Liu , Ruiji Fu , Zhongyuan Wang , Bing Qin

Despite recent advances in large-scale text-to-image generative models, manipulating real images with these models remains a challenging problem. The main limitations of existing editing methods are that they either fail to perform with…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Vadim Titov , Madina Khalmatova , Alexandra Ivanova , Dmitry Vetrov , Aibek Alanov

Existing open-source datasets for arbitrary-instruction image editing remain suboptimal, while a plug-and-play editing module compatible with community-prevalent generative models is notably absent. In this paper, we first introduce the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Jian Ma , Xujie Zhu , Zihao Pan , Qirong Peng , Xu Guo , Chen Chen , Haonan Lu

Introducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose FreeEdit, a novel approach for achieving such reference-based…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Runze He , Kai Ma , Linjiang Huang , Shaofei Huang , Jialin Gao , Xiaoming Wei , Jiao Dai , Jizhong Han , Si Liu

We present a lighting-aware image editing pipeline that, given a portrait image and a text prompt, performs single image relighting. Our model modifies the lighting and color of both the foreground and background to align with the provided…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Junuk Cha , Mengwei Ren , Krishna Kumar Singh , He Zhang , Yannick Hold-Geoffroy , Seunghyun Yoon , HyunJoon Jung , Jae Shin Yoon , Seungryul Baek

Instruction-guided 3D editing is a rapidly emerging field with the potential to broaden access to 3D content creation. However, existing methods face critical limitations: optimization-based approaches are prohibitively slow, while…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Weiwei Cai , Shuangkang Fang , Weicai Ye , Xin Dong , Yunhan Yang , Xuanyang Zhang , Wei Cheng , Yanpei Cao , Gang Yu , Tao Chen

Text-conditioned image editing has emerged as a powerful tool for editing images. However, in many situations, language can be ambiguous and ineffective in describing specific image edits. When faced with such challenges, visual prompts can…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Thao Nguyen , Yuheng Li , Utkarsh Ojha , Yong Jae Lee

We address the task of multi-view image editing from sparse input views, where the inputs can be seen as a mix of images capturing the scene from different viewpoints. The goal is to modify the scene according to a textual instruction while…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Daniel Gilo , Or Litany

Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in modeling high-frequency texture details. We introduce VINS-120K, the first large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Zhizhou Chen , Shanyan Guan , Zhanxin Gao , En Ci , Yanhao Ge , Wei Li , Zhenyu Zhang , Jian Yang , Ying Tai

Text-to-image (T2I) generation has achieved remarkable progress in instruction following and aesthetics. However, a persistent challenge is the prevalence of physical artifacts, such as anatomical and structural flaws, which severely…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Jia Wang , Jie Hu , Xiaoqi Ma , Hanghang Ma , Yanbing Zeng , Xiaoming Wei

This work presents the task of modifying images in an image editing program using natural language written commands. We utilize a corpus of over 6000 image edit text requests to alter real world images collected via crowdsourcing. A novel…

Computation and Language · Computer Science 2018-12-05 Jacqueline Brixey , Ramesh Manuvinakurike , Nham Le , Tuan Lai , Walter Chang , Trung Bui

Image editing aims to edit the given synthetic or real image to meet the specific requirements from users. It is widely studied in recent years as a promising and challenging field of Artificial Intelligence Generative Content (AIGC).…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Xincheng Shuai , Henghui Ding , Xingjun Ma , Rongcheng Tu , Yu-Gang Jiang , Dacheng Tao