English
Related papers

Related papers: A fully differentiable model for unsupervised sing…

200 papers

Informed speaker extraction aims to extract a target speech signal from a mixture of sources given prior knowledge about the desired speaker. Recent deep learning-based methods leverage a speaker discriminative model that maps a reference…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-17 Mohamed Elminshawi , Wolfgang Mack , Emanuël A. P. Habets

Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions in pitch, loudness,…

Sound · Computer Science 2022-10-19 Naoya Takahashi , Mayank Kumar , Singh , Yuki Mitsufuji

We propose a method for the blind separation of sounds of musical instruments in audio signals. We describe the individual tones via a parametric model, training a dictionary to capture the relative amplitudes of the harmonics. The model…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-10 Sören Schulze , Johannes Leuschner , Emily J. King

Single-channel speech separation is a crucial task for enhancing speech recognition systems in multi-speaker environments. This paper investigates the robustness of state-of-the-art Neural Network models in scenarios where the pitch…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-23 Bunlong Lay , Sebastian Zaczek , Kristina Tesch , Timo Gerkmann

General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-24 Bowen Shi , Andros Tjandra , John Hoffman , Helin Wang , Yi-Chiao Wu , Luya Gao , Julius Richter , Matt Le , Apoorv Vyas , Sanyuan Chen , Christoph Feichtenhofer , Piotr Dollár , Wei-Ning Hsu , Ann Lee

Music is often experienced as a progression of concurrent streams of notes, or voices. The degree to which this happens depends on the position along a voice-leading continuum, ranging from monophonic, to homophonic, to polyphonic, which…

Sound · Computer Science 2020-11-06 Patrick Gray , Razvan Bunescu

In this paper we propose a method of single-channel speaker-independent multi-speaker speech separation for an unknown number of speakers. As opposed to previous works, in which the number of speakers is assumed to be known in advance and…

Sound · Computer Science 2019-09-04 Naoya Takahashi , Sudarsanam Parthasaarathy , Nabarun Goswami , Yuki Mitsufuji

Discriminative models for source separation have recently been shown to produce impressive results. However, when operating on sources outside of the training set, these models can not perform as well and are cumbersome to update. Classical…

Sound · Computer Science 2019-11-04 Shrikant Venkataramani , Efthymios Tzinis , Paris Smaragdis

Music source separation is the task of separating a mixture of instruments into constituent tracks. Music source separation models are typically trained using only audio data, although additional information can be used to improve the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Eetu Tunturi , David Diaz-Guerra , Archontis Politis , Tuomas Virtanen

We propose a novel unsupervised singing voice detection method which use single-channel Blind Audio Source Separation (BASS) algorithm as a preliminary step. To reach this goal, we investigate three promising BASS approaches which operate…

Sound · Computer Science 2018-05-04 Dominique Fourer , Geoffroy Peeters

This paper presents a high quality singing synthesizer that is able to model a voice with limited available recordings. Based on the sequence-to-sequence singing model, we design a multi-singer framework to leverage all the existing singing…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-19 Jie Wu , Jian Luan

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

We present a new method for separating a mixed audio sequence, in which multiple voices speak simultaneously. The new method employs gated neural networks that are trained to separate the voices at multiple processing steps, while…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-02 Eliya Nachmani , Yossi Adi , Lior Wolf

In this paper, we study whether music source separation can be used as a pre-training strategy for music representation learning, targeted at music classification tasks. To this end, we first pre-train U-Net networks under various music…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-24 Christos Garoufis , Athanasia Zlatintsi , Petros Maragos

We propose an independence-based joint dereverberation and separation method with a neural source model. We introduce a neural network in the framework of time-decorrelation iterative source steering, which is an extension of independent…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-04 Kohei Saijo , Robin Scheibler

Singing voice conversion is to convert a singer's voice to another one's voice without changing singing content. Recent work shows that unsupervised singing voice conversion can be achieved with an autoencoder-based approach [1]. However,…

Sound · Computer Science 2020-02-19 Chengqi Deng , Chengzhu Yu , Heng Lu , Chao Weng , Dong Yu

We introduce a new method for generating text, and in particular song lyrics, based on the speech-like acoustic qualities of a given audio file. We repurpose a vocal source separation algorithm and an acoustic model trained to recognize…

Human-Computer Interaction · Computer Science 2019-12-17 Jon Gillick , David Bamman

Separating an audio scene into isolated sources is a fundamental problem in computer audition, analogous to image segmentation in visual scene analysis. Source separation systems based on deep learning are currently the most successful…

Sound · Computer Science 2018-11-07 Prem Seetharaman , Gordon Wichern , Jonathan Le Roux , Bryan Pardo

Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or cinematic audio source separation (CASS) with a single…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-01 Kohei Saijo , Janek Ebbers , François G. Germain , Gordon Wichern , Jonathan Le Roux

The objective of deep learning methods based on encoder-decoder architectures for music source separation is to approximate either ideal time-frequency masks or spectral representations of the target music source(s). The spectral…

‹ Prev 1 3 4 5 6 7 10 Next ›