English
Related papers

Related papers: RegionE: Adaptive Region-Aware Generation for Effi…

200 papers

Masked image modeling (MIM) has been recognized as a strong self-supervised pre-training approach in the vision domain. However, the mechanism and properties of the learned representations by such a scheme, as well as how to further enhance…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Kevin Zhang , Zhiqiang Shen

Masked Autoencoders (MAEs) learn generalizable representations for image, text, audio, video, etc., by reconstructing masked input data from tokens of the visible data. Current MAE approaches for videos rely on random patch, tube, or…

Computer Vision and Pattern Recognition · Computer Science 2022-11-17 Wele Gedara Chaminda Bandara , Naman Patel , Ali Gholami , Mehdi Nikkhah , Motilal Agrawal , Vishal M. Patel

The malicious misuse and widespread dissemination of AI-generated images pose a significant threat to the authenticity of online information. Current detection methods often struggle to generalize to unseen generative models, and the rapid…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Hanyi Wang , Jun Lan , Yaoyu Kang , Huijia Zhu , Weiqiang Wang , Zhuosheng Zhang , Shilin Wang

Deep convolutional neural networks (CNN) proved to be highly accurate to perform anatomical segmentation of medical images. However, some of the most popular CNN architectures for image segmentation still rely on post-processing strategies…

Image and Video Processing · Electrical Eng. & Systems 2019-06-07 Agostina J. Larrazabal , Cesar Martinez , Enzo Ferrante

Leveraging visual priors from pre-trained text-to-image (T2I) generative models has shown success in dense prediction. However, dense prediction is inherently an image-to-image task, suggesting that image editing models, rather than T2I…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 JiYuan Wang , Chunyu Lin , Lei Sun , Rongying Liu , Lang Nie , Mingxing Li , Kang Liao , Xiangxiang Chu

Diffusion Transformers (DiTs) achieve state-of-the-art results in text-to-image, text-to-video generation, and editing. However, their large model size and the quadratic cost of spatial-temporal attention over multiple denoising steps make…

Machine Learning · Computer Science 2025-09-24 Muhammad Adnan , Nithesh Kurella , Akhil Arunkumar , Prashant J. Nair

Image defocus is inherent in the physics of image formation caused by the optical aberration of lenses, providing plentiful information on image quality. Unfortunately, existing quality enhancement approaches for compressed images neglect…

Image and Video Processing · Electrical Eng. & Systems 2023-03-14 Qunliang Xing , Mai Xu , Xin Deng , Yichen Guo

Large diffusion transformers (DiTs) follow global editing instructions well but consistently leak local edits into unrelated regions, because joint-attention architectures offer no explicit channel telling the network where to apply the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Honghao Cai , Xiangyuan Wang , Yunhao Bai , Haohua Chen , Tianze Zhou , Runqi Wang , Wei Zhu , Yibo Chen , Xu Tang , Yao Hu , Zhen Li

Recent advances in text-to-image (T2I) models have enabled training-free regional image editing by leveraging the generative priors of foundation models. However, existing methods struggle to balance text adherence in edited regions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Weiyan Xie , Han Gao , Didan Deng , Kaican Li , April Hua Liu , Yongxiang Huang , Nevin L. Zhang

Natural Language Image Editing (NLIE) aims to use natural language instructions to edit images. Since novices are inexperienced with image editing techniques, their instructions are often ambiguous and contain high-level abstractions that…

Computation and Language · Computer Science 2020-02-13 Tzu-Hsiang Lin , Alexander Rudnicky , Trung Bui , Doo Soon Kim , Jean Oh

Mainstream high dynamic range imaging techniques typically rely on fusing multiple images captured with different exposure setups (shutter speed and ISO). A good balance between shutter speed and ISO is crucial for achieving high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Tianyi Xu , Fan Zhang , Boxin Shi , Tianfan Xue , Yujin Wang

Arbitrary resolution image generation provides a consistent visual experience across devices, having extensive applications for producers and consumers. Current diffusion models increase computational demand quadratically with resolution,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Tao Han , Wanghan Xu , Junchao Gong , Xiaoyu Yue , Song Guo , Luping Zhou , Lei Bai

Unified multimodal models often struggle with complex synthesis tasks that demand deep reasoning, and typically treat text-to-image generation and image editing as isolated capabilities rather than interconnected reasoning steps. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Dianyi Wang , Chaofan Ma , Feng Han , Size Wu , Wei Song , Yibin Wang , Zhixiong Zhang , Tianhang Wang , Siyuan Wang , Zhongyu Wei , Jiaqi Wang

Instruction-based image editing aims to modify source content according to textual instructions. However, existing methods built upon flow matching often struggle to maintain consistency in non-edited regions due to denoising-induced…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zongqing Li , Zhihui Liu , Yujie Xie , Shansiyuan Wu , Hongshen Lv , Songzhi Su

All-in-One Image Restoration (AIO-IR) aims to develop a unified model that can handle multiple degradations under complex conditions. However, existing methods often rely on task-specific designs or latent routing strategies, making it hard…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Jingren Liu , Shuning Xu , Qirui Yang , Yun Wang , Xiangyu Chen , Zhong Ji

For autoregressive (AR) modeling of high-resolution images, vector quantization (VQ) represents an image as a sequence of discrete codes. A short sequence length is important for an AR model to reduce its computational costs to consider…

Computer Vision and Pattern Recognition · Computer Science 2022-03-10 Doyup Lee , Chiheon Kim , Saehoon Kim , Minsu Cho , Wook-Shin Han

Image harmonization aims to generate a more realistic appearance of foreground and background for a composite image. Existing methods perform the same harmonization process for the whole foreground. However, the implanted foreground always…

Computer Vision and Pattern Recognition · Computer Science 2022-05-16 Jinlong Peng , Zekun Luo , Liang Liu , Boshen Zhang , Tao Wang , Yabiao Wang , Ying Tai , Chengjie Wang , Weiyao Lin

Given a partial differential equation (PDE), goal-oriented error estimation allows us to understand how errors in a diagnostic quantity of interest (QoI), or goal, occur and accumulate in a numerical approximation, for example using the…

Machine Learning · Computer Science 2022-07-25 Joseph G. Wallwork , Jingyi Lu , Mingrui Zhang , Matthew D. Piggott

Low level image restoration is an integral component of modern artificial intelligence (AI) driven camera pipelines. Most of these frameworks are based on deep neural networks which present a massive computational overhead on resource…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Avisek Lahiri , Sourav Bairagya , Sutanu Bera , Siddhant Haldar , Prabir Kumar Biswas

We introduce Autoregressive Retrieval Augmentation (AR-RAG), a novel paradigm that enhances image generation by autoregressively incorporating knearest neighbor retrievals at the patch level. Unlike prior methods that perform a single,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Jingyuan Qi , Zhiyang Xu , Qifan Wang , Lifu Huang