中文
相关论文

相关论文: MIGE: Mutually Enhanced Multimodal Instruction-Bas…

200 篇论文

Existing text-to-image diffusion models primarily generate images from text prompts. However, the inherent conciseness of textual descriptions poses challenges in faithfully synthesizing images with intricate details, such as specific…

计算机视觉与模式识别 · 计算机科学 2024-06-07 Wei Li , Xue Xu , Jiachen Liu , Xinyan Xiao

The field of video generation has expanded significantly in recent years, with controllable and compositional video generation garnering considerable interest. Most methods rely on leveraging annotations such as text, objects' bounding…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Aram Davtyan , Sepehr Sameni , Björn Ommer , Paolo Favaro

Ensuring precise multimodal alignment between diffusion-generated images and input prompts has been a long-standing challenge. Earlier works finetune diffusion weight using high-quality preference data, which tends to be limited and…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Jiayi Guo , Chuanhao Yan , Xingqian Xu , Yulin Wang , Kai Wang , Gao Huang , Humphrey Shi

Existing AI-driven video creation systems typically treat script drafting and key-shot design as two disjoint tasks: the former relies on large language models, while the latter depends on image generation models. We argue that these two…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Jiaxu Zhang , Tianshu Hu , Yuan Zhang , Zenan Li , Linjie Luo , Guosheng Lin , Xin Chen

This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Zehuan Huang , Yuan-Chen Guo , Xingqiao An , Yunhan Yang , Yangguang Li , Zi-Xin Zou , Ding Liang , Xihui Liu , Yan-Pei Cao , Lu Sheng

Recent advances in instruction-based image editing have shown remarkable progress. However, existing methods remain limited to relatively simple editing operations, hindering real-world applications that require complex and compositional…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Xuehai Bai , Xiaoling Gu , Akide Liu , Hangjie Yuan , YiFan Zhang , Jack Ma

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: the model, the tasks,…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Chi Zhang , Jiepeng Wang , Youming Wang , Yuanzhi Liang , Xiaoyan Yang , Zuoxin Li , Haibin Huang , Xuelong Li

Personalizing image generation and editing is particularly challenging when we only have a few images of the subject, or even a single image. A common approach to personalization is concept learning, which can integrate the subject into…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Yair Shpitzer , Gal Chechik , Idan Schwartz

While modern diffusion models excel at generating high-quality and diverse images, they still struggle with high-fidelity compositional and multimodal control, particularly when users simultaneously specify text prompts, subject references,…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Yusuf Dalva , Guocheng Gordon Qian , Maya Goldenberg , Tsai-Shien Chen , Kfir Aberman , Sergey Tulyakov , Pinar Yanardag , Kuan-Chieh Jackson Wang

This paper presents SPIE: a novel approach for semantic and structural post-training of instruction-based image editing diffusion models, addressing key challenges in alignment with user prompts and consistency with input images. We…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Elior Benarous , Yilun Du , Heng Yang

Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Lorenzo Baraldi , Davide Bucciarelli , Federico Betti , Marcella Cornia , Lorenzo Baraldi , Nicu Sebe , Rita Cucchiara

This is the technique report for the winning solution of the CVPR2024 GenAI Media Generation Challenge Workshop's Instruction-guided Image Editing track. Instruction-guided image editing has been largely studied in recent years. The most…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Xuan Ju , Junhao Zhuang , Zhaoyang Zhang , Yuxuan Bian , Qiang Xu , Ying Shan

We present a Multi-Instance Generation (MIG) task, simultaneously generating multiple instances with diverse controls in one image. Given a set of predefined coordinates and their corresponding descriptions, the task is to ensure that…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Dewei Zhou , You Li , Fan Ma , Xiaoting Zhang , Yi Yang

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Xuan Ju , Tianyu Wang , Yuqian Zhou , He Zhang , Qing Liu , Nanxuan Zhao , Zhifei Zhang , Yijun Li , Yuanhao Cai , Shaoteng Liu , Daniil Pakhomov , Zhe Lin , Soo Ye Kim , Qiang Xu

Empowering models to dynamically accomplish tasks specified through natural language instructions represents a promising path toward more capable and general artificial intelligence. In this work, we introduce InstructSeq, an…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Rongyao Fang , Shilin Yan , Zhaoyang Huang , Jingqiu Zhou , Hao Tian , Jifeng Dai , Hongsheng Li

A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-modal tasks. In this paper, we introduce MMGen, a unified…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Jiepeng Wang , Zhaoqing Wang , Hao Pan , Yuan Liu , Dongdong Yu , Changhu Wang , Wenping Wang

Recent advances in image editing have been driven by the development of denoising diffusion models, marking a significant leap forward in this field. Despite these advances, the generalization capabilities of recent image editing approaches…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Zichong Meng , Changdi Yang , Jun Liu , Hao Tang , Pu Zhao , Yanzhi Wang

The recently developed discrete diffusion models perform extraordinarily well in the text-to-image task, showing significant promise for handling the multi-modality signals. In this work, we harness these traits and present a unified…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Minghui Hu , Chuanxia Zheng , Heliang Zheng , Tat-Jen Cham , Chaoyue Wang , Zuopeng Yang , Dacheng Tao , Ponnuthurai N. Suganthan

Vanilla image completion approaches exhibit sensitivity to large missing regions, attributed to the limited availability of reference information for plausible generation. To mitigate this, existing methods incorporate the extra cue as a…

计算机视觉与模式识别 · 计算机科学 2023-11-22 Yongsheng Yu , Hao Wang , Tiejian Luo , Heng Fan , Libo Zhang

Research in Image Generation has recently made significant progress, particularly boosted by the introduction of Vision-Language models which are able to produce high-quality visual content based on textual inputs. Despite ongoing…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Federico Betti , Jacopo Staiano , Lorenzo Baraldi , Lorenzo Baraldi , Rita Cucchiara , Nicu Sebe