English
Related papers

Related papers: MicAugment: One-shot Microphone Style Transfer

200 papers

Transfer learning is a crucial concept within deep learning that allows artificial neural networks to benefit from a large pre-training data basis when confronted with a task of limited data. Despite its ubiquitous use and clear benefits,…

Machine Learning · Computer Science 2026-05-20 Manuel Milling , Andreas Triantafyllopoulos , Alexander Gebhard , Simon Rampp , Björn W. Schuller

Zero-shot voice conversion aims to transfer the voice of a source speaker to that of a speaker unseen during training, while preserving the content information. Although various methods have been proposed to reconstruct speaker information…

Sound · Computer Science 2024-08-22 Anastasia Avdeeva , Aleksei Gusev

Models for audio generation are typically trained on hours of recordings. Here, we illustrate that capturing the essence of an audio source is typically possible from as little as a few tens of seconds from a single training signal.…

Sound · Computer Science 2021-10-27 Gal Greshler , Tamar Rott Shaham , Tomer Michaeli

The recent surge in popularity of diffusion models for image generation has brought new attention to the potential of these models in other areas of media generation. One area that has yet to be fully explored is the application of…

Sound · Computer Science 2023-02-01 Flavio Schneider

Many hearables contain an in-ear microphone, which may be used to capture the own voice of its user in noisy environments. Since the in-ear microphone mostly records body-conducted speech due to ear canal occlusion, it suffers from…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-25 Mattes Ohlenbusch , Christian Rollwage , Simon Doclo

The subjective evaluation of music generation techniques has been mostly done with questionnaire-based listening tests while ignoring the perspectives from music composition, arrangement, and soundtrack editing. In this paper, we propose an…

Sound · Computer Science 2021-10-26 Wei-Tsung Lu , Meng-Hsuan Wu , Yuh-Ming Chiu , Li Su

Device-guided music transfer adapts playback across unseen devices for users who lack them. Existing methods mainly focus on modifying the timbre, rhythm, harmony, or instrumentation to mimic genres or artists, overlooking the diverse…

Sound · Computer Science 2025-11-24 Manh Pham Hung , Changshuo Hu , Ting Dang , Dong Ma

Diffusion models have shown remarkable progress in text-to-audio generation. However, text-guided audio editing remains in its early stages. This task focuses on modifying the target content within an audio signal while preserving the rest,…

Sound · Computer Science 2026-04-17 Liting Gao , Yi Yuan , Yaru Chen , Yuelan Cheng , Zhenbo Li , Juan Wen , Shubin Zhang , Wenwu Wang

We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider include…

Sound · Computer Science 2023-12-06 Mark Levy , Bruno Di Giorgi , Floris Weers , Angelos Katharopoulos , Tom Nickson

A recently published method for audio style transfer has shown how to extend the process of image style transfer to audio. This method synthesizes audio "content" and "style" independently using the magnitudes of a short time Fourier…

Sound · Computer Science 2017-12-01 Parag K. Mital

Data augmentation methods usually apply the same augmentation (or a mix of them) to all the training samples. For example, to perturb data with noise, the noise is sampled from a Normal distribution with a fixed standard deviation, for all…

Despite rapid progress in the voice style transfer (VST) field, recent zero-shot VST systems still lack the ability to transfer the voice style of a novel speaker. In this paper, we present HierVST, a hierarchical adaptive end-to-end…

Sound · Computer Science 2023-08-01 Sang-Hoon Lee , Ha-Yeong Choi , Hyung-Seok Oh , Seong-Whan Lee

Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps. While long-range dependencies are difficult to model directly in the time domain, we show that they can…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-05 Sean Vasquez , Mike Lewis

While recent automatic speech recognition systems achieve remarkable performance when large amounts of adequate, high quality annotated speech data is used for training, the same systems often only achieve an unsatisfactory result for tasks…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-19 Michael Gref , Oliver Walter , Christoph Schmidt , Sven Behnke , Joachim Köhler

Most people who have tried to learn a foreign language would have experienced difficulties understanding or speaking with a native speaker's accent. For native speakers, understanding or speaking a new accent is likewise a difficult task.…

Sound · Computer Science 2023-10-17 Mumin Jin , Prashant Serai , Jilong Wu , Andros Tjandra , Vimal Manohar , Qing He

The absence of large labeled datasets remains a significant challenge in many application areas of deep learning. Researchers and practitioners typically resort to transfer learning and data augmentation to alleviate this issue. We study…

Sound · Computer Science 2022-11-01 Paul Primus , Gerhard Widmer

Previous studies on music style transfer have mainly focused on one-to-one style conversion, which is relatively limited. When considering the conversion between multiple styles, previous methods required designing multiple modes to…

Sound · Computer Science 2024-04-24 Hong Huang , Yuyi Wang , Luyao Li , Jun Lin

Recently, voice conversion (VC) without parallel data has been successfully adapted to multi-target scenario in which a single model is trained to convert the input voice to many different speakers. However, such model suffers from the…

Machine Learning · Computer Science 2019-08-23 Ju-chieh Chou , Cheng-chieh Yeh , Hung-yi Lee

We propose a model to estimate the fundamental frequency in monophonic audio, often referred to as pitch estimation. We acknowledge the fact that obtaining ground truth annotations at the required temporal and frequency resolution is a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-07 Beat Gfeller , Christian Frank , Dominik Roblek , Matt Sharifi , Marco Tagliasacchi , Mihajlo Velimirović

Reusing recorded sounds (sampling) is a key component in Electronic Music Production (EMP), which has been present since its early days and is at the core of genres like hip-hop or jungle. Commercial and non-commercial services allow users…

Sound · Computer Science 2019-07-22 António Ramires , Xavier Serra