中文
相关论文

相关论文: Open-Source Image Editing Models Are Zero-Shot Vis…

200 篇论文

Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained. While Uncrewed Aerial Vehicles (UAVs) offer scalable data collection, the…

We propose VINO, the first zero-shot, training-free video editing method conditioned on both image and text. Our approach introduces $\rho$-start sampling and dilated dual masking to construct structured noise maps that enable coherent and…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Saemee Choi , Sohyun Jeong , Hyojin Jang , Jaegul Choo , Jinhee Kim

Diffusion models are capable of generating impressive images conditioned on text descriptions, and extensions of these models allow users to edit images at a relatively coarse scale. However, the ability to precisely edit the layout,…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Daniel Geng , Andrew Owens

Current diffusion-based video editing primarily focuses on local editing (\textit{e.g.,} object/background editing) or global style editing by utilizing various dense correspondences. However, these methods often fail to accurately edit the…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Xiangpeng Yang , Linchao Zhu , Hehe Fan , Yi Yang

Text-guided image editing is widely needed in daily life, ranging from personal use to professional applications such as Photoshop. However, existing methods are either zero-shot or trained on an automatically synthesized dataset, which…

计算机视觉与模式识别 · 计算机科学 2024-05-17 Kai Zhang , Lingbo Mo , Wenhu Chen , Huan Sun , Yu Su

High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Xin Gu , Ming Li , Libo Zhang , Fan Chen , Longyin Wen , Tiejian Luo , Sijie Zhu

Recently, large pretrained models (e.g., BERT, StyleGAN, CLIP) have shown great knowledge transfer and generalization capability on various downstream tasks within their domains. Inspired by these efforts, in this paper we propose a unified…

计算机视觉与模式识别 · 计算机科学 2021-12-02 Jing Shi , Ning Xu , Haitian Zheng , Alex Smith , Jiebo Luo , Chenliang Xu

Recent monocular 3D shape reconstruction methods have shown promising zero-shot results on object-segmented images without any occlusions. However, their effectiveness is significantly compromised in real-world conditions, due to imperfect…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Junhyeong Cho , Kim Youwang , Hunmin Yang , Tae-Hyun Oh

This retrospective study evaluated five VLMs (Qwen2.5, Phi-4, Gemma3, Llama3.2, and Mistral3.1) using the MedFMC dataset. This dataset includes 22,349 images from 7,461 patients encompassing chest radiography (19 disease multi-label…

图像与视频处理 · 电气工程与系统科学 2025-08-05 Gustav Müller-Franzes , Debora Jutz , Jakob Nikolas Kather , Christiane Kuhl , Sven Nebelung , Daniel Truhn

Instruction-based image editing aims to modify source content according to textual instructions. However, existing methods built upon flow matching often struggle to maintain consistency in non-edited regions due to denoising-induced…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zongqing Li , Zhihui Liu , Yujie Xie , Shansiyuan Wu , Hongshen Lv , Songzhi Su

Recent advances in generative modeling enable image editing assistants that follow natural language instructions without additional user input. Their supervised training requires millions of triplets (original image, instruction, edited…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Maksim Kuprashevich , Grigorii Alekseenko , Irina Tolstykh , Georgii Fedorov , Bulat Suleimanov , Vladimir Dokholyan , Aleksandr Gordeev

Currently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of vision language models (VLMs). However, they still face challenges in three key areas: 1)…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Jun Zhou , Jiahao Li , Zunnan Xu , Hanhui Li , Yiji Cheng , Fa-Ting Hong , Qin Lin , Qinglin Lu , Xiaodan Liang

State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is needed to specify any…

We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite recent progress, existing models still struggle with ultra-long…

Recent advances in image editing models have shown remarkable progress. A common architectural design couples a multimodal large language model (MLLM) encoder with a diffusion decoder, as seen in systems such as Step1X-Edit and…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Fukun Yin , Shiyu Liu , Yucheng Han , Zhibo Wang , Peng Xing , Rui Wang , Wei Cheng , Yingming Wang , Aojie Li , Zixin Yin , Pengtao Chen , Xiangyu Zhang , Daxin Jiang , Xianfang Zeng , Gang Yu

Adapting pretrained diffusion-based generative models for text-driven image editing with negligible tuning overhead has demonstrated remarkable potential. A classical adaptation paradigm, as followed by these methods, first infers the…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Jiahuan Wang , Yuxin Chen , Jun Yu , Guangming Lu , Wenjie Pei

Text-guided image editing, a pivotal task in modern multimedia content creation, has seen remarkable progress with training-free methods that eliminate the need for additional optimization. Despite recent progress, existing methods are…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Jinhao Shen , Haoqian Du , Xulu Zhang , Xiao-Yong Wei , Qing Li

In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Yan-Bo Lin , Kevin Lin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Chung-Ching Lin , Xiaofei Wang , Gedas Bertasius , Lijuan Wang

Large text-to-image diffusion models have achieved remarkable success in generating diverse, high-quality images. Additionally, these models have been successfully leveraged to edit input images by just changing the text prompt. But when…

计算机视觉与模式识别 · 计算机科学 2023-08-11 Anant Khandelwal

Research in vision-language models has seen rapid developments off-late, enabling natural language-based interfaces for image generation and manipulation. Many existing text guided manipulation techniques are restricted to specific classes…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Paramanand Chandramouli , Kanchana Vaishnavi Gandikota