English
Related papers

Related papers: MIGE: Mutually Enhanced Multimodal Instruction-Bas…

200 papers

This paper introduces a tuning-free method for both object insertion and subject-driven generation. The task involves composing an object, given multiple views, into a scene specified by either an image or text. Existing methods struggle to…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Daniel Winter , Asaf Shul , Matan Cohen , Dana Berman , Yael Pritch , Alex Rav-Acha , Yedid Hoshen

Diffusion-model-based text-guided image generation has recently made astounding progress, producing fascinating results in open-domain image manipulation tasks. Few models, however, currently have complete zero-shot capabilities for both…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Sijia Li , Chen Chen , Haonan Lu

Building on the success of diffusion models in image generation and editing, video editing has recently gained substantial attention. However, maintaining temporal consistency and motion alignment still remains challenging. To address these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Yi Huang , Wei Xiong , He Zhang , Chaoqi Chen , Jianzhuang Liu , Mingfu Yan , Shifeng Chen

Human matting is a foundation task in image and video processing, where human foreground pixels are extracted from the input. Prior works either improve the accuracy by additional guidance or improve the temporal consistency of a single…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Chuong Huynh , Seoung Wug Oh , Abhinav Shrivastava , Joon-Young Lee

Generative depth estimation methods leverage the rich visual priors stored in pre-trained text-to-image diffusion models, demonstrating astonishing zero-shot capability. However, parameter updates during training lead to catastrophic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Hongkai Lin , Dingkang Liang , Mingyang Du , Xin Zhou , Xiang Bai

Masked image generation (MIG) has demonstrated remarkable efficiency and high-fidelity images by enabling parallel token prediction. Existing methods typically rely solely on the model itself to learn semantic dependencies among visual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Guotao Liang , Baoquan Zhang , Zhiyuan Wen , Zihao Han , Yunming Ye

We consider the problem of editing 3D objects and scenes based on open-ended language instructions. A common approach to this problem is to use a 2D image generator or editor to guide the 3D editing process, obviating the need for 3D data.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Minghao Chen , Iro Laina , Andrea Vedaldi

Pre-trained video models learn powerful priors for generating high-quality, temporally coherent content. While these models excel at temporal coherence, their dynamics are often constrained by the continuous nature of their training data.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zhoujie Fu , Xianfang Zeng , Jinghong Lan , Xinyao Liao , Cheng Chen , Junyi Chen , Jiacheng Wei , Wei Cheng , Shiyu Liu , Yunuo Chen , Gang Yu , Guosheng Lin

Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geometric consistency. However, existing methods typically rely on fragmented geometric…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Hong Jiang , Wensong Song , Zongxing Yang , Ruijie Quan , Yi Yang

In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Shaojun E , Yuchen Yang , Jiaheng Wu , Yan Zhang , Tiejun Zhao , Ziyan Chen

Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In contrast, human communication is inherently multimodal, where visual instructions such as…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Huanyu Zhang , Xuehai Bai , Chengzu Li , Chen Liang , Haochen Tian , Haodong Li , Ruichuan An , Yifan Zhang , Anna Korhonen , Zhang Zhang , Liang Wang , Tieniu Tan

We present aMUSEd, an open-source, lightweight masked image model (MIM) for text-to-image generation based on MUSE. With 10 percent of MUSE's parameters, aMUSEd is focused on fast image generation. We believe MIM is under-explored compared…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Suraj Patil , William Berman , Robin Rombach , Patrick von Platen

Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task training. Due to inherent task conflicts, such strategy requires complex multi-stage…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Dian Zheng , Manyuan Zhang , Hongyu Li , Hongbo Liu , Kai Zou , Kaituo Feng , Hongsheng Li

Digital art synthesis is receiving increasing attention in the multimedia community because of engaging the public with art effectively. Current digital art synthesis methods usually use single-modality inputs as guidance, thereby limiting…

Computer Vision and Pattern Recognition · Computer Science 2022-09-29 Nisha Huang , Fan Tang , Weiming Dong , Changsheng Xu

Development of multimodal interactive systems is hindered by the lack of rich, multimodal (text, images) conversational data, which is needed in large quantities for LLMs. Previous approaches augment textual dialogues with retrieved images,…

Computation and Language · Computer Science 2024-10-04 Hossein Aboutalebi , Hwanjun Song , Yusheng Xie , Arshit Gupta , Justin Sun , Hang Su , Igor Shalyminov , Nikolaos Pappas , Siffi Singh , Saab Mansour

Video composition is the core task of video editing. Although image composition based on diffusion models has been highly successful, it is not straightforward to extend the achievement to video object composition tasks, which not only…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Wei Wang , Yaosen Chen , Yuegen Liu , Qi Yuan , Shubin Yang , Yanru Zhang

Class-incremental learning (CIL) requires deep learning models to continuously acquire new knowledge from streaming data while preserving previously learned information. Recently, CIL based on pre-trained models (PTMs) has achieved…

Machine Learning · Computer Science 2025-06-16 Linjie Li , Zhenyu Wu , Yang Ji

Diffusion models have achieved significant success in image and video generation. This motivates a growing interest in video editing tasks, where videos are edited according to provided text descriptions. However, most existing approaches…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Zhen Xing , Qi Dai , Zihao Zhang , Hui Zhang , Han Hu , Zuxuan Wu , Yu-Gang Jiang

Existing multimodal generative models fall short as qualified design copilots, as they often struggle to generate imaginative outputs once instructions are less detailed or lack the ability to maintain consistency with the provided…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Zhipeng Huang , Shaobin Zhuang , Canmiao Fu , Binxin Yang , Ying Zhang , Chong Sun , Zhizheng Zhang , Yali Wang , Chen Li , Zheng-Jun Zha

Generating or editing images directly from Neural signals has immense potential at the intersection of neuroscience, vision, and Brain-computer interaction. In this paper, We present Uni-Neur2Img, a unified framework for neural…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Xiyue Bai , Ronghao Yu , Jia Xiu , Pengfei Zhou , Jie Xia , Peng Ji