中文
相关论文

相关论文: Complex ratio masking for singing voice separation

200 篇论文

Music source separation (MSS) aims to extract 'vocals', 'drums', 'bass' and 'other' tracks from a piece of mixed music. While deep learning methods have shown impressive results, there is a trend toward larger models. In our paper, we…

音频与语音处理 · 电气工程与系统科学 2024-03-20 Junyu Chen , Susmitha Vekkot , Pancham Shukla

Singing Voice Separation (SVS) tries to separate singing voice from a given mixed musical signal. Recently, many U-Net-based models have been proposed for the SVS task, but there were no existing works that evaluate and compare various…

音频与语音处理 · 电气工程与系统科学 2020-10-09 Woosung Choi , Minseok Kim , Jaehwa Chung , Daewon Lee , Soonyoung Jung

Facing the diversity and growth of the musical field nowadays, the search for precise songs becomes more and more complex. The identity of the singer facilitates this search. In this project, we focus on the problem of identifying the…

Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions in pitch, loudness,…

声音 · 计算机科学 2022-10-19 Naoya Takahashi , Mayank Kumar , Singh , Yuki Mitsufuji

Voice conversion (VC) aims to modify the speaker's identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling the speaker's identity from linguistic information is…

音频与语音处理 · 电气工程与系统科学 2024-09-19 Philip H. Lee , Ismail Rasim Ulgen , Berrak Sisman

Neural vocoders have recently advanced waveform generation, yielding natural and expressive audio. Among these approaches, iSTFT-based vocoders have recently gained attention. They predict a complex-valued spectrogram and then synthesize…

声音 · 计算机科学 2026-03-13 Hyung-Seok Oh , Deok-Hyeon Cho , Seung-Bin Kim , Seong-Whan Lee

Music source separation (MSS) aims to separate a music recording into multiple musically distinct stems, such as vocals, bass, drums, and more. Recently, deep learning approaches such as convolutional neural networks (CNNs) and recurrent…

声音 · 计算机科学 2023-09-12 Wei-Tsung Lu , Ju-Chiang Wang , Qiuqiang Kong , Yun-Ning Hung

Time-domain speech enhancement (SE) has recently been intensively investigated. Among recent works, DEMUCS introduces multi-resolution STFT loss to enhance performance. However, some resolutions used for STFT contain non-stationary signals,…

声音 · 计算机科学 2023-03-28 Hao Shi , Masato Mimura , Longbiao Wang , Jianwu Dang , Tatsuya Kawahara

We propose an independence-based joint dereverberation and separation method with a neural source model. We introduce a neural network in the framework of time-decorrelation iterative source steering, which is an extension of independent…

音频与语音处理 · 电气工程与系统科学 2022-04-04 Kohei Saijo , Robin Scheibler

This paper introduces a quantum-inspired denoising framework that integrates the Quantum Fourier Transform (QFT) into classical audio enhancement pipelines. Unlike conventional Fast Fourier Transform (FFT) based methods, QFT provides a…

声音 · 计算机科学 2025-09-08 Rajeshwar Tripathi , Sahil Tomar , Sandeep Kumar , Monika Aggarwal

Speech separation involves extracting an individual speaker's voice from a multi-speaker audio signal. The increasing complexity of real-world environments, where multiple speakers might converse simultaneously, underscores the importance…

音频与语音处理 · 电气工程与系统科学 2024-01-09 Renana Opochinsky , Mordehay Moradi , Sharon Gannot

In this work, we propose a novel consistency-preserving loss function for recovering the phase information in the context of phase reconstruction (PR) and speech enhancement (SE). Different from conventional techniques that directly…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Pin-Jui Ku , Chun-Wei Ho , Hao Yen , Sabato Marco Siniscalchi , Chin-Hui Lee

We study the use of the Wave-U-Net architecture for speech enhancement, a model introduced by Stoller et al for the separation of music vocals and accompaniment. This end-to-end learning method for audio source separation operates directly…

声音 · 计算机科学 2018-11-29 Craig Macartney , Tillman Weyde

In this paper, we introduce a neural network-based method for regional speech separation using a microphone array. This approach leverages novel spatial cues to extract the sound source not only from specified direction but also within…

声音 · 计算机科学 2025-08-12 Yiheng Jiang , Haoxu Wang , Yafeng Chen , Gang Qiao , Biao Tian

This work proposes a multichannel speech separation method with narrow-band Conformer (named NBC). The network is trained to learn to automatically exploit narrow-band speech separation information, such as spatial vector clustering of…

声音 · 计算机科学 2022-07-04 Changsheng Quan , Xiaofei Li

We propose RemixIT, a simple and novel self-supervised training method for speech enhancement. The proposed method is based on a continuously self-training scheme that overcomes limitations from previous studies including assumptions for…

声音 · 计算机科学 2022-11-14 Efthymios Tzinis , Yossi Adi , Vamsi K. Ithapu , Buye Xu , Anurag Kumar

The phase vocoder (PV) is a widely spread technique for processing audio signals. It employs a short-time Fourier transform (STFT) analysis-modify-synthesis loop and is typically used for time-scaling of signals by means of using different…

声音 · 计算机科学 2022-02-16 Zdenek Prusa , Nicki Holighaus

Cochlear implant (CI) users have considerable difficulty in understanding speech in reverberant listening environments. Time-frequency (T-F) masking is a common technique that aims to improve speech intelligibility by multiplying…

音频与语音处理 · 电气工程与系统科学 2021-06-01 Kevin M. Chu , Leslie M. Collins , Boyla O. Mainsah

Speech separation has been extensively studied to deal with the cocktail party problem in recent years. All related approaches can be divided into two categories: time-frequency domain methods and time domain methods. In addition, some…

音频与语音处理 · 电气工程与系统科学 2022-03-31 Fan-Lin Wang , Yu-Huai Peng , Hung-Shin Lee , Hsin-Min Wang

Deep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However,…

机器学习 · 统计学 2017-11-30 Yi Luo , Zhuo Chen , John R. Hershey , Jonathan Le Roux , Nima Mesgarani
‹ 上一页 1 8 9 10 下一页 ›