English
Related papers

Related papers: Music Style Transfer with Time-Varying Inversion o…

200 papers

Given a style-reference image as the additional image condition, text-to-image diffusion models have demonstrated impressive capabilities in generating images that possess the content of text prompts while adopting the visual style of the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-12 Lin Zhu , Xinbing Wang , Chenghu Zhou , Qinying Gu , Nanyang Ye

Content-preserving style transfer, generating stylized outputs based on content and style references, remains a significant challenge for Diffusion Transformers (DiTs) due to the inherent entanglement of content and style features in their…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Shiwen Zhang , Xiaoyan Yang , Bojia Zi , Haibin Huang , Chi Zhang , Xuelong Li

Diffusion models for text-to-image generation, known for their efficiency, accessibility, and quality, have gained popularity. While inference with these systems on consumer-grade GPUs is increasingly feasible, training from scratch…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Bram de Wilde , Anindo Saha , Maarten de Rooij , Henkjan Huisman , Geert Litjens

Synthesizing visually impressive images that seamlessly align both text prompts and specific artistic styles remains a significant challenge in Text-to-Image (T2I) diffusion models. This paper introduces StyleBlend, a method designed to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Zichong Chen , Shijin Wang , Yang Zhou

Multitrack music transcription aims to transcribe a music audio input into the musical notes of multiple instruments simultaneously. It is a very challenging task that typically requires a more complex model to achieve satisfactory result.…

Sound · Computer Science 2023-06-21 Wei-Tsung Lu , Ju-Chiang Wang , Yun-Ning Hung

Synthetic data offers a promising solution for mitigating data scarcity and demographic bias in mental health analysis, yet existing approaches largely rely on pretrained large language models (LLMs), which may suffer from limited output…

Computation and Language · Computer Science 2026-01-21 Saad Mankarious , Aya Zirikly

Diffusion models have recently shown the ability to generate high-quality images. However, controlling its generation process still poses challenges. The image style transfer task is one of those challenges that transfers the visual…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Kento Masui , Mayu Otani , Masahiro Nomura , Hideki Nakayama

We explore the oscillatory behavior observed in inversion methods applied to large-scale text-to-image diffusion models, with a focus on the "Flux" model. By employing a fixed-point-inspired iterative approach to invert real-world images,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yan Zheng , Zhenxiao Liang , Xiaoyan Cong , Lanqing guo , Yuehao Wang , Peihao Wang , Zhangyang Wang

Multimodal and multi-domain stylization are two important problems in the field of image style transfer. Currently, there are few methods that can perform both multimodal and multi-domain stylization simultaneously. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2020-06-03 Minxuan Lin , Fan Tang , Weiming Dong , Xiao Li , Chongyang Ma , Changsheng Xu

Recent years have witnessed significant advancements in text-guided style transfer, primarily attributed to innovations in diffusion models. These models excel in conditional guidance, utilizing text or images to direct the sampling…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Nisha Huang , Kaer Huang , Yifan Pu , Jiangshan Wang , Jie Guo , Yiqiang Yan , Xiu Li , Tong-Yee Lee

This paper explores a variety of models for frame-based music transcription, with an emphasis on the methods needed to reach state-of-the-art on human recordings. The translation-invariant network discussed in this paper, which combines a…

Machine Learning · Statistics 2017-11-15 John Thickstun , Zaid Harchaoui , Dean Foster , Sham M. Kakade

The dominant approach to unsupervised "style transfer" in text is based on the idea of learning a latent representation, which is independent of the attributes specifying its "style". In this paper, we show that this condition is not…

Computation and Language · Computer Science 2019-09-23 Sandeep Subramanian , Guillaume Lample , Eric Michael Smith , Ludovic Denoyer , Marc'Aurelio Ranzato , Y-Lan Boureau

Device-guided music transfer adapts playback across unseen devices for users who lack them. Existing methods mainly focus on modifying the timbre, rhythm, harmony, or instrumentation to mimic genres or artists, overlooking the diverse…

Sound · Computer Science 2025-11-24 Manh Pham Hung , Changshuo Hu , Ting Dang , Dong Ma

Photorealistic style transfer is the task of synthesizing a realistic-looking image when adapting the content from one image to appear in the style of another image. Modern models commonly embed a transformation that fuses features…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Tai-Yin Chiu , Danna Gurari

We introduce Noise2Music, where a series of diffusion models is trained to generate high-quality 30-second music clips from text prompts. Two types of diffusion models, a generator model, which generates an intermediate representation…

This paper introduces a deep-learning approach to photographic style transfer that handles a large variety of image content while faithfully transferring the reference style. Our approach builds upon the recent work on painterly transfer…

Computer Vision and Pattern Recognition · Computer Science 2017-04-12 Fujun Luan , Sylvain Paris , Eli Shechtman , Kavita Bala

We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider include…

Sound · Computer Science 2023-12-06 Mark Levy , Bruno Di Giorgi , Floris Weers , Angelos Katharopoulos , Tom Nickson

Style transfer is a technique for combining two images based on the activations and feature statistics in a deep learning neural network architecture. This paper studies the analogous task in the audio domain and takes a critical look at…

Sound · Computer Science 2020-08-10 M. Huzaifah , L. Wyse

Text-conditioned style transfer enables users to communicate their desired artistic styles through text descriptions, offering a new and expressive means of achieving stylization. In this work, we evaluate the text-conditioned image editing…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Silky Singh , Surgan Jandial , Simra Shahid , Abhinav Java

Text style transfer involves rewriting the content of a source sentence in a target style. Despite there being a number of style tasks with available data, there has been limited systematic discussion of how text style datasets relate to…

Computation and Language · Computer Science 2021-08-19 Stephanie Schoch , Wanyu Du , Yangfeng Ji