English
Related papers

Related papers: Repurposing Image Diffusion Models for Training-Fr…

200 papers

We present Style Matching Score (SMS), a novel optimization method for image stylization with diffusion models. Balancing effective style transfer with content preservation is a long-standing challenge. Unlike existing efforts, our method…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Yuxin Jiang , Liming Jiang , Shuai Yang , Jia-Wei Liu , Ivor Tsang , Mike Zheng Shou

A significant research effort is focused on exploiting the amazing capacities of pretrained diffusion models for the editing of images.They either finetune the model, or invert the image in the latent space of the pretrained model. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Senmao Li , Joost van de Weijer , Taihang Hu , Fahad Shahbaz Khan , Qibin Hou , Yaxing Wang , Jian Yang , Ming-Ming Cheng

Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-17 Siyuan Hou , Shansong Liu , Ruibin Yuan , Wei Xue , Ying Shan , Mangsuo Zhao , Chao Zhang

Text style transfer is usually performed using attributes that can take a handful of discrete values (e.g., positive to negative reviews). In this work, we introduce an architecture that can leverage pre-trained consistent continuous…

Computation and Language · Computer Science 2019-11-12 Eric Michael Smith , Diana Gonzalez-Rico , Emily Dinan , Y-Lan Boureau

With the advance of diffusion models, various personalized image generation methods have been proposed. However, almost all existing work only focuses on either subject-driven or style-driven personalization. Meanwhile, state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Youcan Xu , Zhen Wang , Jun Xiao , Wei Liu , Long Chen

While GAN-based models have been successful in image stylization tasks, they often struggle with structure preservation while stylizing a wide range of input images. Recently, diffusion models have been adopted for image stylization but…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Hansam Cho , Jonghyun Lee , Seunggyu Chang , Yonghyun Jeong

We present Stylos, a single-forward 3D Gaussian framework for 3D style transfer that operates on unposed content, from a single image to a multi-view collection, conditioned on a separate reference style image. Stylos synthesizes a stylized…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hanzhou Liu , Jia Huang , Mi Lu , Srikanth Saripalli , Peng Jiang

In this paper, we show that, a good style representation is crucial and sufficient for generalized style transfer without test-time tuning. We achieve this through constructing a style-aware encoder and a well-organized style dataset called…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Junyao Gao , Yanchen Liu , Yanan Sun , Yinhao Tang , Yanhong Zeng , Kai Chen , Cairong Zhao

Text-to-music generation technology is progressing rapidly, creating new opportunities for musical composition and editing. However, existing music editing methods often fail to preserve the source music's temporal structure, including…

Sound · Computer Science 2025-11-19 Yi Yang , Haowen Li , Tianxiang Li , Boyu Cao , Xiaohan Zhang , Liqun Chen , Qi Liu

Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the inherent entanglement of content and style features in their…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Shiwen Zhang , Xiaoyan Yang , Bojia Zi , Haibin Huang , Chi Zhang , Xuelong Li

This paper creates a novel method of deep neural style transfer by generating style images from freeform user text input. The language model and style transfer model form a seamless pipeline that can create output images with similar losses…

Computer Vision and Pattern Recognition · Computer Science 2022-12-15 Tejas Santanam , Mengyang Liu , Jiangyue Yu , Zhaodong Yang

Video style transfer aims to alter the style of a video while preserving its content. Previous methods often struggle with content leakage and style misalignment, particularly when using image-driven approaches that aim to transfer precise…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Jiang Lin , Zili Yi

Large-scale noisy web image-text datasets have been proven to be efficient for learning robust vision-language models. However, when transferring them to the task of video retrieval, models still need to be fine-tuned on hand-curated paired…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Nina Shvetsova , Anna Kukleva , Bernt Schiele , Hilde Kuehne

Stylized text-to-image generation focuses on creating images from textual descriptions while adhering to a style specified by a few reference images. However, subtle style variations within different reference images can hinder the model…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Xing Cui , Zekun Li , Pei Pei Li , Huaibo Huang , Xuannan Liu , Zhaofeng He

Consumer-grade music recordings such as those captured by mobile devices typically contain distortions in the form of background noise, reverb, and microphone-induced EQ. This paper presents a deep learning approach to enhance low-quality…

Sound · Computer Science 2022-04-29 Nikhil Kandpal , Oriol Nieto , Zeyu Jin

The goal of image style transfer is to render an image guided by a style reference while maintaining the original content. Existing image-guided methods rely on specific style reference images, restricting their wider application and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Yuexing Han , Liheng Ruan , Bing Wang

Style transfer is a technique for combining two images based on the activations and feature statistics in a deep learning neural network architecture. This paper studies the analogous task in the audio domain and takes a critical look at…

Sound · Computer Science 2020-08-10 M. Huzaifah , L. Wyse

Deep generative models are now able to synthesize high-quality audio signals, shifting the critical aspect in their development from audio quality to control capabilities. Although text-to-music generation is getting largely adopted by the…

Sound · Computer Science 2024-08-02 Nils Demerlé , Philippe Esling , Guillaume Doras , David Genova

Despite the impressive results of arbitrary image-guided style transfer methods, text-driven image stylization has recently been proposed for transferring a natural image into a stylized one according to textual descriptions of the target…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Nisha Huang , Yuxin Zhang , Fan Tang , Chongyang Ma , Haibin Huang , Yong Zhang , Weiming Dong , Changsheng Xu

Cross-modality image segmentation aims to segment the target modalities using a method designed in the source modality. Deep generative models can translate the target modality images into the source modality, thus enabling cross-modality…

Image and Video Processing · Electrical Eng. & Systems 2024-04-11 Zihao Wang , Yingyu Yang , Yuzhou Chen , Tingting Yuan , Maxime Sermesant , Herve Delingette , Ona Wu