English
Related papers

Related papers: Qwen-Image-Layered: Towards Inherent Editability v…

200 papers

Generative diffusion models trained on large-scale datasets have achieved remarkable progress in image synthesis. In favor of their ability to supplement missing details and generate aesthetically pleasing contents, recent works have…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Junhao Cheng , Wei-Ting Chen , Xi Lu , Ming-Hsuan Yang

Many image-to-image translation problems are ambiguous, as a single input image may correspond to multiple possible outputs. In this work, we aim to model a \emph{distribution} of possible outputs in a conditional generative modeling…

Computer Vision and Pattern Recognition · Computer Science 2018-10-25 Jun-Yan Zhu , Richard Zhang , Deepak Pathak , Trevor Darrell , Alexei A. Efros , Oliver Wang , Eli Shechtman

It is a challenging task to restore images from their variants with combined distortions. In the existing works, a promising strategy is to apply parallel "operations" to handle different types of distortion. However, in the feature fusion…

Computer Vision and Pattern Recognition · Computer Science 2020-01-30 Zihao Huang , Chao Li , Feng Duan , Qibin Zhao

While text-to-image models have achieved impressive capabilities in image generation and editing, their application across various modalities often necessitates training separate models. Inspired by existing method of single image editing…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Gihyun Kwon , Jangho Park , Jong Chul Ye

Disentangling factors of variation has become a very challenging problem on representation learning. Existing algorithms suffer from many limitations, such as unpredictable disentangling factors, poor quality of generated images from…

Computer Vision and Pattern Recognition · Computer Science 2018-03-29 Taihong Xiao , Jiapeng Hong , Jinwen Ma

Existing 3D-aware facial generation methods face a dilemma in quality versus editability: they either generate editable results in low resolution or high-quality ones with no editing flexibility. In this work, we propose a new approach that…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Jingxiang Sun , Xuan Wang , Yichun Shi , Lizhen Wang , Jue Wang , Yebin Liu

The task of image-to-multi-view generation refers to generating novel views of an instance from a single image. Recent methods achieve this by extending text-to-image latent diffusion models to multi-view version, which contains an VAE…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Zhenggang Tang , Peiye Zhuang , Chaoyang Wang , Aliaksandr Siarohin , Yash Kant , Alexander Schwing , Sergey Tulyakov , Hsin-Ying Lee

Large-scale text-to-image models enable a wide range of image editing techniques, using text prompts or even spatial controls. However, applying these editing methods to multi-view images depicting a single scene leads to 3D-inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2024-02-23 Or Patashnik , Rinon Gal , Daniel Cohen-Or , Jun-Yan Zhu , Fernando De la Torre

Semantic image synthesis is a process for generating photorealistic images from a single semantic mask. To enrich the diversity of multimodal image synthesis, previous methods have controlled the global appearance of an output image by…

Computer Vision and Pattern Recognition · Computer Science 2021-06-30 Yuki Endo , Yoshihiro Kanamori

Diffusion probabilistic models (DPMs) have shown remarkable results on various image synthesis tasks such as text-to-image generation and image inpainting. However, compared to other generative methods like VAEs and GANs, DPMs lack a…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Yipeng Leng , Qiangjuan Huang , Zhiyuan Wang , Yangyang Liu , Haoyu Zhang

Face editing methods, essential for tasks like virtual avatars, digital human synthesis and identity preservation, have traditionally been built upon GAN-based techniques, while recent focus has shifted to diffusion-based models due to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Mengting Wei , Tuomas Varanka , Yante Li , Xingxun Jiang , Huai-Qian Khor , Guoying Zhao

Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks which couple powerful vision-language model (VLM) with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Yuxin Song , Wenkai Dong , Shizun Wang , Qi Zhang , Song Xue , Tao Yuan , Hu Yang , Haocheng Feng , Hang Zhou , Xinyan Xiao , Jingdong Wang

Evaluating diffusion-based image-editing models is a crucial task in the field of Generative AI. Specifically, it is imperative to assess their capacity to execute diverse editing tasks while preserving the image content and realism. While…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Stefan Stefanache , Lluís Pastor Pérez , Julen Costa Watanabe , Ernesto Sanchez Tejedor , Thomas Hofmann , Enis Simsar

Multimodal editing large models have demonstrated powerful editing capabilities across diverse tasks. However, a persistent and long-standing limitation is the decline in facial identity (ID) consistency during realistic portrait editing.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Yuran Dong , Hang Dai , Mang Ye

Generative image editing has recently witnessed extremely fast-paced growth. Some works use high-level conditioning such as text, while others use low-level conditioning. Nevertheless, most of them lack fine-grained control over the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Vidit Goel , Elia Peruzzo , Yifan Jiang , Dejia Xu , Xingqian Xu , Nicu Sebe , Trevor Darrell , Zhangyang Wang , Humphrey Shi

Generative Artificial Intelligence (GenAI) features such as image editing, object removal, and prompt-guided image transformation are increasingly integrated into mobile applications. However, deploying Large Vision Models (LVMs) for such…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Sowmya Vajrala , Aakash Parmar , Prasanna R , Sravanth Kodavanti , Manjunath Arveti , Srinivas Soumitri Miriyala , Ashok Senapati

We present Qwen-Image-VAE-2.0, a suite of high-compression Variational Autoencoders (VAEs) that achieve significant advances in both reconstruction fidelity and diffusability. To address the reconstruction bottlenecks of high compression,…

Diffusion models dominate image editing, yet their global denoising mechanism entangles edited regions with surrounding context, causing modifications to propagate into areas that should remain intact. We propose a fundamentally different…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Wei Chow , Linfeng Li , Xian Sun , Lingdong Kong , Zefeng Li , Qi Xu , Hang Song , Tian Ye , Xian Wang , Jinbin Bai , Shilin Xu , Xiangtai Li , Junting Pan , Shaoteng Liu , Ran Zhou , Tianshu Yang , Songhua Liu

Federated learning is a machine learning paradigm that enables decentralized clients to collaboratively learn a shared model while keeping all the training data local. While considerable research has focused on federated image generation,…

Machine Learning · Computer Science 2025-05-06 Chen Hu , Hanchi Ren , Jingjing Deng , Xianghua Xie , Xiaoke Ma

Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Qihang Zhang , Shuangfei Zhai , Miguel Angel Bautista , Kevin Miao , Alexander Toshev , Joshua Susskind , Jiatao Gu