English
Related papers

Related papers: DreamOmni2: Multimodal Instruction-based Editing a…

200 papers

Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text prompts for instruction-based editing and generation, but language often fails to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Bin Xia , Bohao Peng , Jiyang Liu , Sitong Wu , Jingyao Li , Junjia Huang , Xu Zhao , Yitong Wang , Ruihang Chu , Bei Yu , Jiaya Jia

Editing images via instruction provides a natural way to generate interactive content, but it is a big challenge due to the higher requirement of scene understanding and generation. Prior work utilizes a chain of large language models,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Liya Ji , Chenyang Qi , Qifeng Chen

Instruction-based image editing holds immense potential for a variety of applications, as it enables users to perform any editing operation using a natural language instruction. However, current models in this domain often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Shelly Sheynin , Adam Polyak , Uriel Singer , Yuval Kirstain , Amit Zohar , Oron Ashual , Devi Parikh , Yaniv Taigman

Subject-driven image generation aims at generating images containing customized subjects, which has recently drawn enormous attention from the research community. However, the previous works cannot precisely control the background and…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Tianle Li , Max Ku , Cong Wei , Wenhu Chen

Currently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, streamline deployment, and foster synergistic benefits across different tasks. However, in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Bin Xia , Yuechen Zhang , Jingyao Li , Chengyao Wang , Yitong Wang , Xinglong Wu , Bei Yu , Jiaya Jia

Current instruction-based editing methods, such as InstructPix2Pix, often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this, this paper…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Yuzhou Huang , Liangbin Xie , Xintao Wang , Ziyang Yuan , Xiaodong Cun , Yixiao Ge , Jiantao Zhou , Chao Dong , Rui Huang , Ruimao Zhang , Ying Shan

This paper presents instruct-imagen, a model that tackles heterogeneous image generation tasks and generalizes across unseen tasks. We introduce *multi-modal instruction* for image generation, a task representation articulating a range of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 Hexiang Hu , Kelvin C. K. Chan , Yu-Chuan Su , Wenhu Chen , Yandong Li , Kihyuk Sohn , Yang Zhao , Xue Ben , Boqing Gong , William Cohen , Ming-Wei Chang , Xuhui Jia

In this paper, we focus on the task of instruction-based image editing. Previous works like InstructPix2Pix, InstructDiffusion, and SmartEdit have explored end-to-end editing. However, two limitations still remain: First, existing datasets…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Yingjing Xu , Jie Kong , Jiazhi Wang , Xiao Pan , Bo Lin , Qiang Liu

Current text-driven image editing methods typically follow one of two directions: relying on large-scale, high-quality editing pair datasets to improve editing precision and diversity, or exploring alternative dataset-free techniques.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Chenrui Ma , Xi Xiao , Tianyang Wang , Yanning Shen

The ability to provide fine-grained control for generating and editing visual imagery has profound implications for computer vision and its applications. Previous works have explored extending controllability in two directions: instruction…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Shufan Li , Harkanwar Singh , Aditya Grover

Recent advances in diffusion-based video generation have substantially improved visual fidelity and temporal coherence. However, most existing approaches remain task-specific and rely primarily on textual instructions, limiting their…

Given an original image, image editing aims to generate an image that align with the provided instruction. The challenges are to accept multimodal inputs as instructions and a scarcity of high-quality training data, including crucial…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Zhen Han , Chaojie Mao , Zeyinzi Jiang , Yulin Pan , Jingfeng Zhang

Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to accurately execute complex user instructions, as they are trained on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Qifan Yu , Wei Chow , Zhongqi Yue , Kaihang Pan , Yang Wu , Xiaoyang Wan , Juncheng Li , Siliang Tang , Hanwang Zhang , Yueting Zhuang

Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them separately, struggling with limited high-quality data and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Xueyun Tian , Wei Li , Bingbing Xu , Yige Yuan , Yuanzhuo Wang , Huawei Shen

This paper introduces a novel dataset construction pipeline that samples pairs of frames from videos and uses multimodal large language models (MLLMs) to generate editing instructions for training instruction-based image manipulation…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Mingdeng Cao , Xuaner Zhang , Yinqiang Zheng , Zhihao Xia

Instruction-guided image editing methods have demonstrated significant potential by training diffusion models on automatically synthesized or manually annotated image editing pairs. However, these methods remain far from practical,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Cong Wei , Zheyang Xiong , Weiming Ren , Xinrun Du , Ge Zhang , Wenhu Chen

$360^{\circ}$ omnidirectional images (ODIs) have gained considerable attention recently, and are widely used in various virtual reality (VR) and augmented reality (AR) applications. However, capturing such images is expensive and requires…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Liu Yang , Huiyu Duan , Yucheng Zhu , Xiaohong Liu , Lu Liu , Zitong Xu , Guangji Ma , Xiongkuo Min , Guangtao Zhai , Patrick Le Callet

Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions. However, these models often struggle with complex instructions involving combinatorial editing operations or inter-step…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zilai Zeng , Mingdeng Cao , Zijie Li , Xiaochen Lian , Yichun Shi , Peihao Zhu , Chen Sun , Peng Wang

As information exists in various modalities in real world, effective interaction and fusion among multimodal information plays a key role for the creation and perception of multimodal data in computer vision and deep learning research. With…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Fangneng Zhan , Yingchen Yu , Rongliang Wu , Jiahui Zhang , Shijian Lu , Lingjie Liu , Adam Kortylewski , Christian Theobalt , Eric Xing

Instruction-based image editing focuses on equipping a generative model with the capacity to adhere to human-written instructions for editing images. Current approaches typically comprehend explicit and specific instructions. However, they…

Computer Vision and Pattern Recognition · Computer Science 2024-06-03 Ying Jin , Pengyang Ling , Xiaoyi Dong , Pan Zhang , Jiaqi Wang , Dahua Lin
‹ Prev 1 2 3 10 Next ›