English
Related papers

Related papers: EditVerse: Unifying Image and Video Editing and Ge…

200 papers

Instruction-guided image editing has achieved remarkable progress, yet current models still face challenges with complex instructions and often require multiple samples to produce a desired result. Reinforcement Learning (RL) offers a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Xin Luo , Jiahao Wang , Chenyuan Wu , Shitao Xiao , Xiyan Jiang , Defu Lian , Jiajun Zhang , Dong Liu , Zheng liu

Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks which couple powerful vision-language model (VLM) with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Yuxin Song , Wenkai Dong , Shizun Wang , Qi Zhang , Song Xue , Tao Yuan , Hu Yang , Haocheng Feng , Hang Zhou , Xinyan Xiao , Jingdong Wang

Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We challenge this paradigm by asking whether all modalities, including textual descriptions,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Junchao Yi , Rui Zhao , Jiahao Tang , Weixian Lei , Linjie Li , Qisheng Su , Zhengyuan Yang , Lijuan Wang , Xiaofeng Zhu , Alex Jinpeng Wang

Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to accurately execute complex user instructions, as they are trained on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Qifan Yu , Wei Chow , Zhongqi Yue , Kaihang Pan , Yang Wu , Xiaoyang Wan , Juncheng Li , Siliang Tang , Hanwang Zhang , Yueting Zhuang

The quality and diversity of instruction-based image editing datasets are continuously increasing, yet large-scale, high-quality datasets for instruction-based video editing remain scarce. To address this gap, we introduce OpenVE-3M, an…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Haoyang He , Jie Wang , Jiangning Zhang , Zhucun Xue , Xingyuan Bu , Qiangpeng Yang , Shilei Wen , Lei Xie

Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them separately, struggling with limited high-quality data and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Xueyun Tian , Wei Li , Bingbing Xu , Yige Yuan , Yuanzhuo Wang , Huawei Shen

We introduce UniVerse-1, a unified, Veo-3-like model capable of simultaneously generating coordinated audio and video. To enhance training efficiency, we bypass training from scratch and instead employ a stitching of experts (SoE)…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Duomin Wang , Wei Zuo , Aojie Li , Ling-Hao Chen , Xinyao Liao , Deyu Zhou , Zixin Yin , Xili Dai , Daxin Jiang , Gang Yu

Visual-prompt-guided edit transfer aims to learn image transformations directly from example pairs, offering more precise and controllable editing than purely text-driven approaches. However, existing diffusion transformer-based methods…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Lan Chen , Qi Mao , Yiren Song , Yuchao Gu , Siwei Ma

Although a video is effectively a sequence of images, visual perception systems typically model images and videos separately, thus failing to exploit the correlation and the synergy provided by these two media. While a few prior research…

Computer Vision and Pattern Recognition · Computer Science 2019-06-13 Yufei Wang , Du Tran , Lorenzo Torresani

A plethora of text-guided image editing methods have recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models such as Imagen and Stable Diffusion. A standardized evaluation protocol,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Samyadeep Basu , Mehrdad Saberi , Shweta Bhardwaj , Atoosa Malemir Chegini , Daniela Massiceti , Maziar Sanjabi , Shell Xu Hu , Soheil Feizi

Unified multimodal models target joint understanding, reasoning, and generation, but current image editing benchmarks are largely confined to natural images and shallow commonsense reasoning, offering limited assessment of this capability…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Mingxin Liu , Ziqian Fan , Zhaokai Wang , Leyao Gu , Zirun Zhu , Yiguo He , Yuchen Yang , Changyao Tian , Xiangyu Zhao , Ning Liao , Shaofeng Zhang , Qibing Ren , Zhihang Zhong , Xuanhe Zhou , Junchi Yan , Xue Yang

Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow-based generation due to text-visual token imbalance and the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Jiabin Luo , Junhui Lin , Zeyu Zhang , Biao Wu , Meng Fang , Ling Chen , Hao Tang

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

Unifying diverse image generation tasks within a single framework remains a fundamental challenge in visual generation. While large language models (LLMs) achieve unification through task-agnostic data and generation, existing visual…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yijing Lin , Mengqi Huang , Shuhan Zhuang , Zhendong Mao

Adapting pretrained diffusion-based generative models for text-driven image editing with negligible tuning overhead has demonstrated remarkable potential. A classical adaptation paradigm, as followed by these methods, first infers the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Jiahuan Wang , Yuxin Chen , Jun Yu , Guangming Lu , Wenjie Pei

Diffusion Transformer has demonstrated powerful capability and scalability in generating high-quality images and videos. Further pursuing the unification of generation and editing tasks has yielded significant progress in the domain of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Zeyinzi Jiang , Zhen Han , Chaojie Mao , Jingfeng Zhang , Yulin Pan , Yu Liu

Recent advances in image editing have been driven by the development of denoising diffusion models, marking a significant leap forward in this field. Despite these advances, the generalization capabilities of recent image editing approaches…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Zichong Meng , Changdi Yang , Jun Liu , Hao Tang , Pu Zhao , Yanzhi Wang

The burgeoning field of Artificial Intelligence Generated Content (AIGC) is witnessing rapid advancements, particularly in video generation. This paper introduces AIGCBench, a pioneering comprehensive and scalable benchmark designed to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Fanda Fan , Chunjie Luo , Wanling Gao , Jianfeng Zhan

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Junyi Chen , Tong He , Zhoujie Fu , Pengfei Wan , Kun Gai , Weicai Ye

Following the advancements in text-guided image generation technology exemplified by Stable Diffusion, video generation is gaining increased attention in the academic community. However, relying solely on text guidance for video generation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Cong Wang , Jiaxi Gu , Panwen Hu , Haoyu Zhao , Yuanfan Guo , Jianhua Han , Hang Xu , Xiaodan Liang