中文
相关论文

相关论文: InstructAny2Pix: Flexible Visual Editing via Multi…

200 篇论文

Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for multi-modal pretraining, which can benefit from modality…

计算机视觉与模式识别 · 计算机科学 2023-05-12 Haiyang Xu , Qinghao Ye , Ming Yan , Yaya Shi , Jiabo Ye , Yuanhong Xu , Chenliang Li , Bin Bi , Qi Qian , Wei Wang , Guohai Xu , Ji Zhang , Songfang Huang , Fei Huang , Jingren Zhou

Large foundation models have recently emerged as a prominent focus of interest, attaining superior performance in widespread scenarios. Due to the scarcity of 3D data, many efforts have been made to adapt pre-trained transformers from…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Yiwen Tang , Ray Zhang , Jiaming Liu , Zoey Guo , Dong Wang , Zhigang Wang , Bin Zhao , Shanghang Zhang , Peng Gao , Hongsheng Li , Xuelong Li

Prompts have been proven to play a crucial role in large language models, and in recent years, vision models have also been using prompts to improve scalability for multiple downstream tasks. In this paper, we focus on adapting prompt…

计算机视觉与模式识别 · 计算机科学 2023-05-02 Zhenxiang Xiao , Yuzhong Chen , Lu Zhang , Junjie Yao , Zihao Wu , Xiaowei Yu , Yi Pan , Lin Zhao , Chong Ma , Xinyu Liu , Wei Liu , Xiang Li , Yixuan Yuan , Dinggang Shen , Dajiang Zhu , Tianming Liu , Xi Jiang

We present PandaGPT, an approach to emPower large lANguage moDels with visual and Auditory instruction-following capabilities. Our pilot experiments show that PandaGPT can perform complex tasks such as detailed image description generation,…

计算与语言 · 计算机科学 2023-05-29 Yixuan Su , Tian Lan , Huayang Li , Jialu Xu , Yan Wang , Deng Cai

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

Despite the advances in text-to-image synthesis, particularly with diffusion models, generating visual instructions that require consistent representation and smooth state transitions of objects across sequential steps remains a formidable…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Quynh Phung , Songwei Ge , Jia-Bin Huang

Image modality is not perfect as it often fails in certain conditions, e.g., night and fast motion. This significantly limits the robustness and versatility of existing multi-modal (i.e., Image+X) semantic segmentation methods when…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Xu Zheng , Yuanhuiyi Lyu , Lin Wang

Image editing affords increased control over the aesthetics and content of generated images. Pre-existing works focus predominantly on text-based instructions to achieve desired image modifications, which limit edit precision and accuracy.…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Bowen Li , Yongxin Yang , Steven McDonagh , Shifeng Zhang , Petru-Daniel Tudosiu , Sarah Parisot

Diffusion models have achieved significant success in image and video generation. This motivates a growing interest in video editing tasks, where videos are edited according to provided text descriptions. However, most existing approaches…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Zhen Xing , Qi Dai , Zihao Zhang , Hui Zhang , Han Hu , Zuxuan Wu , Yu-Gang Jiang

Multimodal Large Language Models (MLLMs) demonstrate remarkable performance across a wide range of domains, with increasing emphasis on enhancing their zero-shot generalization capabilities for unseen tasks across various modalities.…

Recent neural talking radiance field methods have shown great success in photorealistic audio-driven talking face synthesis. In this paper, we propose a novel interactive framework that utilizes human instructions to edit such implicit…

计算机视觉与模式识别 · 计算机科学 2023-08-17 Yuqi Sun , Ruian He , Weimin Tan , Bo Yan

Instruction-based image editing is among the fastest developing areas in generative AI. Over the past year, the field has reached a new level, with dozens of open-source models released alongside highly capable commercial systems. However,…

Recent advancements in diffusion models have significantly facilitated text-guided video editing. However, there is a relative scarcity of research on image-guided video editing, a method that empowers users to edit videos by merely…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Zhi-Lin Huang , Yixuan Liu , Chujun Qin , Zhongdao Wang , Dong Zhou , Dong Li , Emad Barsoum

We propose a generative model that, given a coarsely edited image, synthesizes a photorealistic output that follows the prescribed layout. Our method transfers fine details from the original image and preserve the identity of its parts.…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Hadi Alzayer , Zhihao Xia , Xuaner Zhang , Eli Shechtman , Jia-Bin Huang , Michael Gharbi

In this report, I present an inpainting framework named \textit{ControlFill}, which involves training two distinct prompts: one for generating plausible objects within a designated mask (\textit{creation}) and another for filling the region…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Boseong Jeon

Diffusion-model-based text-guided image generation has recently made astounding progress, producing fascinating results in open-domain image manipulation tasks. Few models, however, currently have complete zero-shot capabilities for both…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Sijia Li , Chen Chen , Haonan Lu

Controlled video generation has seen drastic improvements in recent years. However, editing actions and dynamic events, or inserting contents that should affect the behaviors of other objects in real-world videos, remains a major challenge.…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Vladimir Kulikov , Roni Paiss , Andrey Voynov , Inbar Mosseri , Tali Dekel , Tomer Michaeli

Text-to-image diffusion models have emerged as an evolutionary for producing creative content in image synthesis. Based on the impressive generation abilities of these models, instruction-guided diffusion models can edit images with simple…

密码学与安全 · 计算机科学 2024-08-21 Ruoxi Chen , Haibo Jin , Yixin Liu , Jinyin Chen , Haohan Wang , Lichao Sun

Current machine learning models for vision are often highly specialized and limited to a single modality and task. In contrast, recent large language models exhibit a wide range of capabilities, hinting at a possibility for similarly…

计算机视觉与模式识别 · 计算机科学 2023-12-12 David Mizrahi , Roman Bachmann , Oğuzhan Fatih Kar , Teresa Yeo , Mingfei Gao , Afshin Dehghan , Amir Zamir

Vision-and-language navigation requires an agent to navigate through a real 3D environment following natural language instructions. Despite significant advances, few previous works are able to fully utilize the strong correspondence between…

计算机视觉与模式识别 · 计算机科学 2020-10-06 Yicong Hong , Cristian Rodriguez-Opazo , Qi Wu , Stephen Gould