中文
相关论文

相关论文: Optimisation-Based Multi-Modal Semantic Image Edit…

200 篇论文

Multi-modal visual understanding of images with prompts involves using various visual and textual cues to enhance the semantic understanding of images. This approach combines both vision and language processing to generate more accurate…

计算机视觉与模式识别 · 计算机科学 2023-05-17 Yuzhou Peng

Language has emerged as a natural interface for image editing. In this paper, we introduce a method for region-based image editing driven by textual prompts, without the need for user-provided masks or sketches. Specifically, our approach…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Yuanze Lin , Yi-Wen Chen , Yi-Hsuan Tsai , Lu Jiang , Ming-Hsuan Yang

Diffusion model based language-guided image editing has achieved great success recently. However, existing state-of-the-art diffusion models struggle with rendering correct text and text style during generation. To tackle this problem, we…

计算机视觉与模式识别 · 计算机科学 2023-10-19 Haoxing Chen , Zhuoer Xu , Zhangxuan Gu , Jun Lan , Xing Zheng , Yaohui Li , Changhua Meng , Huijia Zhu , Weiqiang Wang

Instruction-based image editing has made a great process in using natural human language to manipulate the visual content of images. However, existing models are limited by the quality of the dataset and cannot accurately localize editing…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Tiancheng Li , Jinxiu Liu , Huajun Chen , Qi Liu

Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained spatial manipulations, motivating a dedicated assessment…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Yicheng Xiao , Wenhu Zhang , Lin Song , Yukang Chen , Wenbo Li , Nan Jiang , Tianhe Ren , Haokun Lin , Wei Huang , Haoyang Huang , Xiu Li , Nan Duan , Xiaojuan Qi

Visual aesthetic assessment has been an active research field for decades. Although latest methods have achieved promising performance on benchmark datasets, they typically rely on a large number of manual annotations including both…

计算机视觉与模式识别 · 计算机科学 2019-12-04 Kekai Sheng , Weiming Dong , Menglei Chai , Guohui Wang , Peng Zhou , Feiyue Huang , Bao-Gang Hu , Rongrong Ji , Chongyang Ma

Text-guided diffusion models have significantly advanced image editing, enabling highly realistic and local modifications based on textual prompts. While these developments expand creative possibilities, their malicious use poses…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Valentina Bazyleva , Nicolo Bonettini , Gaurav Bharaj

Research in vision-language models has seen rapid developments off-late, enabling natural language-based interfaces for image generation and manipulation. Many existing text guided manipulation techniques are restricted to specific classes…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Paramanand Chandramouli , Kanchana Vaishnavi Gandikota

With the rapid advancement of commercial multi-modal models, image editing has garnered significant attention due to its widespread applicability in daily life. Despite impressive progress, existing image editing systems, particularly…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yiran Zhao , Yaoqi Ye , Xiang Liu , Michael Qizhe Shieh , Trung Bui

Text-guided image editing aims to modify specific regions according to the target prompt while preserving the identity of the source image. Recent methods exploit explicit binary masks to constrain editing, but hard mask boundaries…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Yongwen Lai , Chaoqun Wang , Shaobo Min

Despite recent advances in large-scale text-to-image generative models, manipulating real images with these models remains a challenging problem. The main limitations of existing editing methods are that they either fail to perform with…

计算机视觉与模式识别 · 计算机科学 2024-09-26 Vadim Titov , Madina Khalmatova , Alexandra Ivanova , Dmitry Vetrov , Aibek Alanov

We address the task of multi-view image editing from sparse input views, where the inputs can be seen as a mix of images capturing the scene from different viewpoints. The goal is to modify the scene according to a textual instruction while…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Daniel Gilo , Or Litany

Instruction-based image editing, which aims to modify the image faithfully according to the instruction while preserving irrelevant content unchanged, has made significant progress. However, there still lacks a comprehensive metric for…

图形学 · 计算机科学 2025-06-18 Zhuoying Li , Zhu Xu , Yuxin Peng , Yang Liu

For a given scene, humans can easily reason for the locations and pose to place objects. Designing a computational model to reason about these affordances poses a significant challenge, mirroring the intuitive reasoning abilities of humans.…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Rishubh Parihar , Harsh Gupta , Sachidanand VS , R. Venkatesh Babu

Precise and flexible image editing remains a fundamental challenge in computer vision. Based on the modified areas, most editing methods can be divided into two main types: global editing and local editing. In this paper, we choose the two…

计算机视觉与模式识别 · 计算机科学 2025-02-27 Ziqi Jiang , Zhen Wang , Long Chen

Current instruction-based editing methods, such as InstructPix2Pix, often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this, this paper…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Yuzhou Huang , Liangbin Xie , Xintao Wang , Ziyang Yuan , Xiaodong Cun , Yixiao Ge , Jiantao Zhou , Chao Dong , Rui Huang , Ruimao Zhang , Ying Shan

Recent advances in image editing have been driven by the development of denoising diffusion models, marking a significant leap forward in this field. Despite these advances, the generalization capabilities of recent image editing approaches…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Zichong Meng , Changdi Yang , Jun Liu , Hao Tang , Pu Zhao , Yanzhi Wang

Natural language offers a highly intuitive interface for image editing. In this paper, we introduce the first solution for performing local (region-based) edits in generic natural images, based on a natural language description along with…

计算机视觉与模式识别 · 计算机科学 2023-03-22 Omri Avrahami , Dani Lischinski , Ohad Fried

Training-free image editing with large diffusion models has become practical, yet faithfully performing complex non-rigid edits (e.g., pose or shape changes) remains highly challenging. We identify a key underlying cause: attention collapse…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Zhuo Chen , Fanyue Wei , Runze Xu , Jingjing Li , Lixin Duan , Angela Yao , Wen Li

Large-scale pre-trained diffusion models empower users to edit images through text guidance. However, existing methods often over-align with target prompts while inadequately preserving source image semantics. Such approaches generate…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Jianda Mao , Kaibo Wang , Yang Xiang , Kani Chen