中文
相关论文

相关论文: Raw Multi-Channel Audio Source Separation using Mu…

200 篇论文

Single-channel speech separation in time domain and frequency domain has been widely studied for voice-driven applications over the past few years. Most of previous works assume known number of speakers in advance, however, which is not…

音频与语音处理 · 电气工程与系统科学 2020-04-02 Yiming Xiao , Haijian Zhang

Separating vocal elements from musical tracks is a longstanding challenge in audio signal processing. This study tackles the distinct separation of vocal components from musical spectrograms. We employ the Short Time Fourier Transform…

声音 · 计算机科学 2024-05-31 Adam Sorrenti

In this paper, we study whether music source separation can be used as a pre-training strategy for music representation learning, targeted at music classification tasks. To this end, we first pre-train U-Net networks under various music…

音频与语音处理 · 电气工程与系统科学 2024-04-24 Christos Garoufis , Athanasia Zlatintsi , Petros Maragos

The objective of this paper is to perform audio-visual sound source separation, i.e.~to separate component audios from a mixture based on the videos of sound sources. Moreover, we aim to pinpoint the source location in the input video…

计算机视觉与模式识别 · 计算机科学 2021-04-20 Lingyu Zhu , Esa Rahtu

In deep neural networks with convolutional layers, each layer typically has fixed-size/single-resolution receptive field (RF). Convolutional layers with a large RF capture global information from the input features, while layers with small…

声音 · 计算机科学 2017-11-01 Emad M. Grais , Hagen Wierstorf , Dominic Ward , Mark D. Plumbley

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main…

声音 · 计算机科学 2020-01-03 Rongzhi Gu , Yuexian Zou

The audio source separation tasks, such as speech enhancement, speech separation, and music source separation, have achieved impressive performance in recent studies. The powerful modeling capabilities of deep neural networks give us hope…

音频与语音处理 · 电气工程与系统科学 2021-07-15 Lu Zhang , Chenxing Li , Feng Deng , Xiaorui Wang

The current paradigm for creating and deploying immersive audio content is based on audio objects, which are composed of an audio track and position metadata. While rendering an object-based production into a multichannel mix is…

声音 · 计算机科学 2021-12-22 Daniel Arteaga , Jordi Pons

Query-based audio source extraction seeks to recover a target source from a mixture conditioned on a query. Existing approaches are largely confined to single-channel audio, leaving the spatial information in multi-channel recordings…

音频与语音处理 · 电气工程与系统科学 2025-10-16 Chenxin Yu , Hao Ma , Xu Li , Xiao-Lei Zhang , Mingjie Shao , Chi Zhang , Xuelong Li

In this paper, we propose a source separation method that is trained by observing the mixtures and the class labels of the sources present in the mixture without any access to isolated sources. Since our method does not require source class…

声音 · 计算机科学 2019-08-06 Ertuğ Karamatlı , Ali Taylan Cemgil , Serap Kırbız

We present a novel source separation model to decompose asingle-channel speech signal into two speech segments belonging to two different speakers. The proposed model is a neural network based on residual blocks, and uses learnt speaker…

声音 · 计算机科学 2019-06-25 Shuo Liu , Gil Keren , Björn Schuller

Sound event detection systems typically consist of two stages: extracting hand-crafted features from the raw audio waveform, and learning a mapping between these features and the target sound events using a classifier. Recently, the focus…

声音 · 计算机科学 2018-05-11 Emre Çakır , Tuomas Virtanen

We introduce a new audio processing technique that increases the sampling rate of signals such as speech or music using deep convolutional neural networks. Our model is trained on pairs of low and high-quality audio examples; at test-time,…

声音 · 计算机科学 2017-08-03 Volodymyr Kuleshov , S. Zayd Enam , Stefano Ermon

In this paper we propose a method for separation of moving sound sources. The method is based on first tracking the sources and then estimation of source spectrograms using multichannel non-negative matrix factorization (NMF) and extracting…

声音 · 计算机科学 2017-10-30 Joonas Nikunen , Aleksandr Diment , Tuomas Virtanen

In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the component sources from…

音频与语音处理 · 电气工程与系统科学 2020-10-27 Scott Wisdom , Efthymios Tzinis , Hakan Erdogan , Ron J. Weiss , Kevin Wilson , John R. Hershey

Audio source separation is often achieved by estimating the magnitude spectrogram of each source, and then applying a phase recovery (or spectrogram inversion) algorithm to retrieve time-domain signals. Typically, spectrogram inversion is…

声音 · 计算机科学 2023-07-03 Paul Magron , Tuomas Virtanen

Traditional methods to tackle many music information retrieval tasks typically follow a two-step architecture: feature engineering followed by a simple learning algorithm. In these "shallow" architectures, feature engineering and learning…

声音 · 计算机科学 2015-11-18 Peter Li , Jiyuan Qian , Tian Wang

Motivated by the fact that characteristics of different sound classes are highly diverse in different temporal scales and hierarchical levels, a novel deep convolutional neural network (CNN) architecture is proposed for the environmental…

声音 · 计算机科学 2018-06-15 Boqing Zhu , Kele Xu , Dezhi Wang , Lilun Zhang , Bo Li , Yuxing Peng

State-of-the-art text-independent speaker verification systems typically use cepstral features or filter bank energies as speech features. Recent studies attempted to extract speaker embeddings directly from raw waveforms and have shown…

音频与语音处理 · 电气工程与系统科学 2021-06-10 Ge Zhu , Fei Jiang , Zhiyao Duan

Recently, significant progress has been made in audio source separation by the application of deep learning techniques. Current methods that combine both audio and visual information use 2D representations such as images to guide the…

声音 · 计算机科学 2021-02-04 Francesc Lluís , Vasileios Chatziioannou , Alex Hofmann