English
Related papers

Related papers: MAST: Mask-Guided Attention Mass Allocation for Tr…

200 papers

Although diffusion models have demonstrated remarkable generative capabilities, existing style transfer techniques often struggle to maintain identity while achieving high-quality stylization. This limitation becomes particularly critical…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Mohammad Ali Rezaei , Helia Hajikazem , Saeed Khanehgir , Mahdi Javanmardi

Fashion attribute editing is a task that aims to convert the semantic attributes of a given fashion image while preserving the irrelevant regions. Previous works typically employ conditional GANs where the generator explicitly learns the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Chaerin Kong , DongHyeon Jeon , Ohjoon Kwon , Nojun Kwak

The rapid expansion of large-scale text-to-image diffusion models has raised growing concerns regarding their potential misuse in creating harmful or misleading content. In this paper, we introduce MACE, a finetuning framework for the task…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Shilin Lu , Zilan Wang , Leyang Li , Yanzhu Liu , Adams Wai-Kin Kong

We present Multiscale Audio Spectrogram Transformer (MAST) for audio classification, which brings the concept of multiscale feature hierarchies to the Audio Spectrogram Transformer (AST). Given an input audio spectrogram, we first patchify…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Sreyan Ghosh , Ashish Seth , S. Umesh , Dinesh Manocha

We present HyperNST; a neural style transfer (NST) technique for the artistic stylization of images, based on Hyper-networks and the StyleGAN2 architecture. Our contribution is a novel method for inducing style transfer parameterized by a…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Dan Ruta , Andrew Gilbert , Saeid Motiian , Baldo Faieta , Zhe Lin , John Collomosse

Recent advances in diffusion models have significantly improved the performance of reference-guided line art colorization. However, existing methods still struggle with region-level color consistency, especially when the reference and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Qianru Qiu , Jiafeng Mao , Kento Masui , Xueting Wang

Diffusion models have become leading approaches for high-fidelity image generation. Recent DiT-based diffusion models, in particular, achieve strong prompt adherence while producing high-quality samples. We propose SHIFT, a simple but…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Nina Konovalova , Andrey Kuznetsov , Aibek Alanov

Despite significant advancements in image generation using advanced generative frameworks, cross-image integration of content and style remains a key challenge. Current generative models, while powerful, frequently depend on vague textual…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Shaoxu Li , Ye Pan

With the rising focus on quadrupeds, a generalized policy capable of handling different robot models and sensor inputs becomes highly beneficial. Although several methods have been proposed to address different morphologies, it remains a…

Robotics · Computer Science 2025-03-13 Dikai Liu , Tianwei Zhang , Jianxiong Yin , Simon See

Even in the era of rapid advances in large models, video understanding remains a highly challenging task. Compared to texts or images, videos commonly contain more information with redundancy, requiring large models to properly allocate…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Shiwen Cao , Zhaoxing Zhang , Junming Jiao , Juyi Qiao , Guowen Song , Rong Shen , Xiangbing Meng

Text-to-image diffusion models have recently become highly capable, yet their behavior in multi-object scenes remains unreliable: models often produce an incorrect number of instances and exhibit semantics leaking across objects. We trace…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Sanghyun Jo , Wooyeol Lee , Ziseok Lee , Kyungsu Kim

Recent advances in generative diffusion models have shown a notable inherent understanding of image style and semantics. In this paper, we leverage the self-attention features from pretrained diffusion networks to transfer the visual…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Yang Zhou , Xu Gao , Zichong Chen , Hui Huang

Recent studies have shown remarkable success in universal style transfer which transfers arbitrary visual styles to content images. However, existing approaches suffer from the aesthetic-unrealistic problem that introduces disharmonious…

Computer Vision and Pattern Recognition · Computer Science 2022-08-30 Zhizhong Wang , Zhanjie Zhang , Lei Zhao , Zhiwen Zuo , Ailin Li , Wei Xing , Dongming Lu

Incremental learning aims to adapt to new sets of categories over time with minimal computational overhead. Prior work often addresses this task by training efficient task-specific adaptors that modify frozen layer weights or features to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Nazia Tasnim , Bryan A. Plummer

Adapting a large language model for multiple-attribute text style transfer via fine-tuning can be challenging due to the significant amount of computational resources and labeled data required for the specific task. In this paper, we…

Computation and Language · Computer Science 2023-05-11 Zhiqiang Hu , Roy Ka-Wei Lee , Nancy F. Chen

To prevent unauthorized use of text in images, Scene Text Removal (STR) has become a crucial task. It focuses on automatically removing text and replacing it with a natural, text-less background while preserving significant details such as…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Sanhita Pathak , Vinay Kaushik , Brejesh Lall

Diffusion-based text-to-image models have rapidly gained popularity for their ability to generate detailed and realistic images from textual descriptions. However, these models often reflect the biases present in their training data,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Hidir Yesiltepe , Kiymet Akdemir , Pinar Yanardag

Text-driven style transfer aims to merge the style of a reference image with content described by a text prompt. Recent advancements in text-to-image models have improved the nuance of style transformations, yet significant challenges…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Mingkun Lei , Xue Song , Beier Zhu , Hao Wang , Chi Zhang

Facial attribute editing and style manipulation are crucial for applications like virtual avatars and photo editing. However, achieving precise control over facial attributes without altering unrelated features is challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Wenmin Huang , Weiqi Luo , Xiaochun Cao , Jiwu Huang

Mixing style transfer automates the generation of a multitrack mix for a given set of tracks by inferring production attributes from a reference song. However, existing systems for mixing style transfer are limited in that they often…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-15 Soumya Sai Vanka , Christian Steinmetz , Jean-Baptiste Rolland , Joshua Reiss , George Fazekas