中文
相关论文

相关论文: Compute and memory efficient universal sound sourc…

200 篇论文

Transformers have recently achieved state-of-the-art performance in speech separation. These models, however, are computationally demanding and require a lot of learnable parameters. This paper explores Transformer-based speech separation…

音频与语音处理 · 电气工程与系统科学 2024-01-17 Luca Della Libera , Cem Subakan , Mirco Ravanelli , Samuele Cornell , Frédéric Lepoutre , François Grondin

The state of the art in music source separation employs neural networks trained in a supervised fashion on multi-track databases to estimate the sources from a given mixture. With only few datasets available, often extensive data…

机器学习 · 计算机科学 2018-04-09 Daniel Stoller , Sebastian Ewert , Simon Dixon

Generative adversarial networks (GANs) and diffusion models have recently achieved state-of-the-art performance in audio super-resolution (ADSR), producing perceptually convincing wideband audio from narrowband inputs. However, existing…

音频与语音处理 · 电气工程与系统科学 2026-01-08 Mikhail Silaev , Konstantinos Drossos , Tuomas Virtanen

Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA)…

声音 · 计算机科学 2024-03-22 Samuel Pegg , Kai Li , Xiaolin Hu

Deep convolutional neural networks are known to specialize in distilling compact and robust prior from a large amount of data. We are interested in applying deep networks in the absence of training dataset. In this paper, we introduce deep…

声音 · 计算机科学 2019-12-24 Yapeng Tian , Chenliang Xu , Dingzeyu Li

Recently, research on audio foundation models has witnessed notable advances, as illustrated by the ever improving results on complex downstream tasks. Subsequently, those pretrained networks have quickly been used for various audio…

声音 · 计算机科学 2025-02-19 David Genova , Philippe Esling , Tom Hurlin

General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed…

In this paper, we present a new single sound source DOA estimation and tracking system based on the well-known SRP-PHAT algorithm and a three-dimensional Convolutional Neural Network. It uses SRP-PHAT power maps as input features of a fully…

音频与语音处理 · 电气工程与系统科学 2020-12-18 David Diaz-Guerra , Antonio Miguel , Jose R. Beltran

Deep learning-based methods have made significant achievements in music source separation. However, obtaining good results while maintaining a low model complexity remains challenging in super wide-band music source separation. Previous…

音频与语音处理 · 电气工程与系统科学 2024-01-25 Weinan Tong , Jiaxu Zhu , Jun Chen , Shiyin Kang , Tao Jiang , Yang Li , Zhiyong Wu , Helen Meng

Recently, direct modeling of raw waveforms using deep neural networks has been widely studied for a number of tasks in audio domains. In speaker verification, however, utilization of raw waveforms is in its preliminary phase, requiring…

音频与语音处理 · 电气工程与系统科学 2019-07-18 Jee-weon Jung , Hee-Soo Heo , Ju-ho Kim , Hye-jin Shim , Ha-Jin Yu

Multiple moving sound source localization in real-world scenarios remains a challenging issue due to interaction between sources, time-varying trajectories, distorted spatial cues, etc. In this work, we propose to use deep learning…

声音 · 计算机科学 2022-02-17 Bing Yang , Hong Liu , Xiaofei Li

Audio source separation is often used as preprocessing of various applications, and one of its ultimate goals is to construct a single versatile model capable of dealing with the varieties of audio signals. Since sampling frequency, one of…

声音 · 计算机科学 2021-05-11 Koichi Saito , Tomohiko Nakamura , Kohei Yatabe , Yuma Koizumi , Hiroshi Saruwatari

A three-stage approach is proposed for speaker counting and speech separation in noisy and reverberant environments. In the spatial feature extraction, a spatial coherence matrix (SCM) is computed using whitened relative transfer functions…

音频与语音处理 · 电气工程与系统科学 2023-08-08 Yicheng Hsu , Mingsian Bai

Music source separation has been a popular topic in signal processing for decades, not only because of its technical difficulty, but also due to its importance to many commercial applications, such as automatic karoake and remixing. In this…

音频与语音处理 · 电气工程与系统科学 2020-03-23 Yuzhou Liu , Balaji Thoshkahna , Ali Milani , Trausti Kristjansson

Speech separation in realistic acoustic environments remains challenging because overlapping speakers, background noise, and reverberation must be resolved simultaneously. Although recent time-frequency (TF) domain models have shown strong…

音频与语音处理 · 电气工程与系统科学 2026-05-15 Ui-Hyeop Shin , Hyung-Min Park

Audio source separation is often achieved by estimating the magnitude spectrogram of each source, and then applying a phase recovery (or spectrogram inversion) algorithm to retrieve time-domain signals. Typically, spectrogram inversion is…

声音 · 计算机科学 2023-07-03 Paul Magron , Tuomas Virtanen

The performance of audio source separation from underdetermined convolutive mixture assuming known mixing filters can be significantly improved by using an analysis sparse prior optimized by a reweighting l1 scheme and a wideband…

声音 · 计算机科学 2015-06-18 Simon Arberet , Pierre Vandergheynst

Transformer based end-to-end modelling approaches with multiple stream inputs have been achieved great success in various automatic speech recognition (ASR) tasks. An important issue associated with such approaches is that the intermediate…

音频与语音处理 · 电气工程与系统科学 2022-07-11 Jin Li , Rongfeng Su , Xurong Xie , Nan Yan , Lan Wang

We address the problem of acoustic source separation in a deep learning framework we call "deep clustering." Rather than directly estimating signals or masking functions, we train a deep network to produce spectrogram embeddings that are…

神经与进化计算 · 计算机科学 2015-08-19 John R. Hershey , Zhuo Chen , Jonathan Le Roux , Shinji Watanabe

This paper describes a versatile method that accelerates multichannel source separation methods based on full-rank spatial modeling. A popular approach to multichannel source separation is to integrate a spatial model with a source model…

声音 · 计算机科学 2019-03-11 Kouhei Sekiguchi , Aditya Arie Nugraha , Yoshiaki Bando , Kazuyoshi Yoshii