中文
相关论文

相关论文: Hybrid Y-Net Architecture for Singing Voice Separa…

200 篇论文

Deep learning-based music source separation has gained a lot of interest in the last decades. Most of the existing methods operate with either spectrograms or waveforms. Spectrogram based models learn suitable masks for separating magnitude…

音频与语音处理 · 电气工程与系统科学 2021-12-10 Chin-Yun Yu , Kin-Wai Cheuk

The goal of this work is to investigate what singing voice separation approaches based on neural networks learn from the data. We examine the mapping functions of neural networks based on the denoising autoencoder (DAE) model that are…

音频与语音处理 · 电气工程与系统科学 2019-10-22 Stylianos Ioannis Mimilakis , Konstantinos Drossos , Estefanía Cano , Gerald Schuller

Music source separation is important for applications such as karaoke and remixing. Much of previous research focuses on estimating short-time Fourier transform (STFT) magnitude and discarding phase information. We observe that, for singing…

音频与语音处理 · 电气工程与系统科学 2020-11-05 Yixuan Zhang , Yuzhou Liu , DeLiang Wang

We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages recent advances of…

音频与语音处理 · 电气工程与系统科学 2020-02-26 Yin-Jyun Luo , Chin-Chen Hsu , Kat Agres , Dorien Herremans

Fully-supervised models for source separation are trained on parallel mixture-source data and are currently state-of-the-art. However, such parallel data is often difficult to obtain, and it is cumbersome to adapt trained models to mixtures…

音频与语音处理 · 电气工程与系统科学 2022-11-30 Ge Zhu , Jordan Darefsky , Fei Jiang , Anton Selitskiy , Zhiyao Duan

Music source separation in the time-frequency domain is commonly achieved by applying a soft or binary mask to the magnitude component of (complex) spectrograms. The phase component is usually not estimated, but instead copied from the…

声音 · 计算机科学 2021-03-25 Andreas Jansson , Rachel M. Bittner , Nicola Montecchio , Tillman Weyde

Recent approaches for music source separation are almost exclusively based on deep neural networks, mostly employing recurrent neural networks (RNNs). Although RNNs are in many cases superior than other types of deep neural networks for…

音频与语音处理 · 电气工程与系统科学 2020-07-08 Pyry Pyykkönen , Styliannos I. Mimilakis , Konstantinos Drossos , Tuomas Virtanen

In the recent years, singing voice separation systems showed increased performance due to the use of supervised training. The design of training datasets is known as a crucial factor in the performance of such systems. We investigate on how…

声音 · 计算机科学 2019-06-07 Laure Prétet , Romain Hennequin , Jimena Royo-Letelier , Andrea Vaglio

Separation of competing speech is a key challenge in signal processing and a feat routinely performed by the human auditory brain. A long standing benchmark of the spectrogram approach to source separation is known as the ideal binary mask.…

声音 · 计算机科学 2015-03-25 Andrew J. R. Simpson

Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker…

音频与语音处理 · 电气工程与系统科学 2019-09-24 Jian Wu , Yong Xu , Shi-Xiong Zhang , Lian-Wu Chen , Meng Yu , Lei Xie , Dong Yu

Extracting direct-path spatial feature is crucial for sound source localization in adverse acoustic environments. This paper proposes the IPDnet, a neural network that estimates direct-path inter-channel phase difference (DP-IPD) of sound…

音频与语音处理 · 电气工程与系统科学 2024-05-14 Yabo Wang , Bing Yang , Xiaofei Li

Recently, audio-visual separation approaches have taken advantage of the natural synchronization between the two modalities to boost audio source separation performance. They extracted high-level semantics from visual inputs as the guidance…

声音 · 计算机科学 2024-07-08 Shentong Mo , Yapeng Tian

Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions in pitch, loudness,…

声音 · 计算机科学 2022-10-19 Naoya Takahashi , Mayank Kumar , Singh , Yuki Mitsufuji

While neural network approaches have made significant strides in resolving classical signal processing problems, it is often the case that hybrid approaches that draw insight from both signal processing and neural networks produce more…

音频与语音处理 · 电气工程与系统科学 2024-02-13 Karim Helwani , Masahito Togami , Paris Smaragdis , Michael M. Goodwin

Deep neural networks with convolutional layers usually process the entire spectrogram of an audio signal with the same time-frequency resolutions, number of filters, and dimensionality reduction scale. According to the constant-Q transform,…

声音 · 计算机科学 2019-10-22 Emad M. Grais , Fei Zhao , Mark D. Plumbley

Recent work has shown that recurrent neural networks can be trained to separate individual speakers in a sound mixture with high fidelity. Here we explore convolutional neural network models as an alternative and show that they achieve…

声音 · 计算机科学 2018-05-29 Shariq Mobin , Brian Cheung , Bruno Olshausen

Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different…

信号处理 · 电气工程与系统科学 2021-02-09 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

This study presents UX-Net, a time-domain audio separation network (TasNet) based on a modified U-Net architecture. The proposed UX-Net works in real-time and handles either single or multi-microphone input. Inspired by the…

音频与语音处理 · 电气工程与系统科学 2022-10-31 Kashyap Patel , Anton Kovalyov , Issa Panahi

Singing voice detection (SVD), to recognize vocal parts in the song, is an essential task in music information retrieval (MIR). The task remains challenging since singing voice varies and intertwines with the accompaniment music, especially…

音频与语音处理 · 电气工程与系统科学 2022-05-09 Yifu Sun , Xulong Zhang , Yi Yu , Xi Chen , Wei Li

Melody extraction in polyphonic musical audio is important for music signal processing. In this paper, we propose a novel streamlined encoder/decoder network that is designed for the task. We make two technical contributions. First, drawing…

音频与语音处理 · 电气工程与系统科学 2019-02-19 Tsung-Han Hsieh , Li Su , Yi-Hsuan Yang