中文
相关论文

相关论文: RegionE: Adaptive Region-Aware Generation for Effi…

200 篇论文

Recently, the editing of neural radiance fields (NeRFs) has gained considerable attention, but most prior works focus on static scenes while research on the appearance editing of dynamic scenes is relatively lacking. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Shangzan Zhang , Sida Peng , Yinji ShenTu , Qing Shuai , Tianrun Chen , Kaicheng Yu , Hujun Bao , Xiaowei Zhou

Latent generative models have emerged as a leading approach for high-quality image synthesis. These models rely on an autoencoder to compress images into a latent space, followed by a generative model to learn the latent distribution. We…

机器学习 · 计算机科学 2025-08-05 Theodoros Kouzelis , Ioannis Kakogeorgiou , Spyros Gidaris , Nikos Komodakis

In this paper, we propose a framework called TrustMAE to address the problem of product defect classification. Instead of relying on defective images that are difficult to collect and laborious to label, our framework can accept datasets…

计算机视觉与模式识别 · 计算机科学 2021-01-01 Daniel Stanley Tan , Yi-Chun Chen , Trista Pei-Chun Chen , Wei-Chao Chen

Text-to-image diffusion models have achieved remarkable progress in recent years. However, training models for high-resolution image generation remains challenging, particularly when training data and computational resources are limited. In…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Ruonan Yu , Songhua Liu , Zhenxiong Tan , Xinchao Wang

This work presents 3DPE, a practical method that can efficiently edit a face image following given prompts, like reference images or text descriptions, in a 3D-aware manner. To this end, a lightweight module is distilled from a 3D portrait…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Qingyan Bai , Zifan Shi , Yinghao Xu , Hao Ouyang , Qiuyu Wang , Ceyuan Yang , Xuan Wang , Gordon Wetzstein , Yujun Shen , Qifeng Chen

Diffeomorphic deformable multi-modal image registration is a challenging task which aims to bring images acquired by different modalities to the same coordinate space and at the same time to preserve the topology and the invertibility of…

图像与视频处理 · 电气工程与系统科学 2022-03-16 Vasiliki Sideri-Lampretsa , Georgios Kaissis , Daniel Rueckert

Text-guided image editing involves modifying a source image based on a language instruction and, typically, requires changes to only small local regions. However, existing approaches generate the entire target image rather than selectively…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Huimin Wu , Xiaojian Ma , Haozhe Zhao , Yanpeng Zhao , Qing Li

We introduce the Self-Evaluating Model (Self-E), a novel, from-scratch training approach for text-to-image generation that supports any-step inference. Self-E learns from data similarly to a Flow Matching model, while simultaneously…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Xin Yu , Xiaojuan Qi , Zhengqi Li , Kai Zhang , Richard Zhang , Zhe Lin , Eli Shechtman , Tianyu Wang , Yotam Nitzan

Denoising diffusion models have enabled high-quality image generation and editing. We present a method to localize the desired edit region implicit in a text instruction. We leverage InstructPix2Pix (IP2P) and identify the discrepancy…

Efficiently modeling spatial-temporal information in videos is crucial for action recognition. To achieve this goal, state-of-the-art methods typically employ the convolution operator and the dense interaction modules such as non-local…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Yuan Tian , Yichao Yan , Guangtao Zhai , Guodong Guo , Zhiyong Gao

While latent diffusion models achieve impressive image editing results, their application to iterative editing of the same image is severely restricted. When trying to apply consecutive edit operations using current models, they accumulate…

图形学 · 计算机科学 2025-04-29 Gal Almog , Ariel Shamir , Ohad Fried

Inspired by the masked language modeling (MLM) in natural language processing tasks, the masked image modeling (MIM) has been recognized as a strong self-supervised pre-training method in computer vision. However, the high random mask ratio…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Zhaowen Li , Yousong Zhu , Zhiyang Chen , Wei Li , Chaoyang Zhao , Rui Zhao , Ming Tang , Jinqiao Wang

Disentangled and interpretable latent representations in generative models typically come at the cost of generation quality. The $\beta$-VAE framework introduces a hyperparameter $\beta$ to balance disentanglement and reconstruction…

机器学习 · 计算机科学 2025-07-10 Anshuk Uppal , Yuhta Takida , Chieh-Hsin Lai , Yuki Mitsufuji

Point-drag-based image editing methods, like DragDiffusion, have attracted significant attention. However, point-drag-based approaches suffer from computational overhead and misinterpretation of user intentions due to the sparsity of…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Jingyi Lu , Xinghui Li , Kai Han

The image-to-image generation task aims to produce controllable images by leveraging conditional inputs and prompt instructions. However, existing methods often train separate control branches for each type of condition, leading to…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Guoqing Zhang , Xingtong Ge , Lu Shi , Xin Zhang , Muqing Xue , Wanru Xu , Yigang Cen , Yidong Li

High-resolution image editing is essential for professional and creative applications, yet existing multimodal diffusion-based editors remain computationally inefficient and constrained to relatively low resolutions. Current approaches…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Yuyao Zhang , Alexander Huang-Menders , Yu-Wing Tai

Generative modeling has evolved to a notable field of machine learning. Deep polynomial neural networks (PNNs) have demonstrated impressive results in unsupervised image generation, where the task is to map an input vector (i.e., noise) to…

机器学习 · 计算机科学 2021-10-29 Grigorios G Chrysos , Markos Georgopoulos , Yannis Panagakis

Multimodal Large Language Models (MLLMs) often struggle to accurately perceive fine-grained visual details, especially when targets are tiny or visually subtle. This challenge can be addressed through semantic-visual information fusion,…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Yuxiang Shen , Hailong Huang , Zhenkun Gao , Xueheng Li , Man Zhou , Chengjun Xie , Haoxuan Che , Xuanhua He , Jie Zhang

Diffusion Transformers (DiTs) achieve state-of-the-art performance in text-to-image synthesis but remain computationally expensive due to the iterative nature of denoising and the quadratic cost of global attention. In this work, we observe…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Bowen Lin , Fanjiang Ye , Yihua Liu , Zhenghui Guo , Boyuan Zhang , Weijian Zheng , Yufan Xu , Tiancheng Xing , Yuke Wang , Chengming Zhang

Automatic detecting anomalous regions in images of objects or textures without priors of the anomalies is challenging, especially when the anomalies appear in very small areas of the images, making difficult-to-detect visual variations,…

计算机视觉与模式识别 · 计算机科学 2022-07-05 Jie Yang , Yong Shi , Zhiquan Qi