中文
相关论文

相关论文: Mask-Weighted Spatial Likelihood Coding for Speake…

200 篇论文

In this work, we present a method for learning interpretable music signal representations directly from waveform signals. Our method can be trained using unsupervised objectives and relies on the denoising auto-encoder model that uses a…

音频与语音处理 · 电气工程与系统科学 2020-07-02 Stylianos I. Mimilakis , Konstantinos Drossos , Gerald Schuller

Speaker diarization systems are challenged by a trade-off between the temporal resolution and the fidelity of the speaker representation. By obtaining a superior temporal resolution with an enhanced accuracy, a multi-scale approach is a way…

音频与语音处理 · 电气工程与系统科学 2022-03-31 Tae Jin Park , Nithin Rao Koluguri , Jagadeesh Balam , Boris Ginsburg

Despite the recent success of deep learning for many speech processing tasks, single-microphone, speaker-independent speech separation remains challenging for two main reasons. The first reason is the arbitrary order of the target and…

声音 · 计算机科学 2018-04-19 Yi Luo , Zhuo Chen , Nima Mesgarani

Speech encodes multiple simultaneous attributes -- linguistic content, speaker identity, dialect, gender --that conventional single-vector embeddings conflate. We present a factor-partitioned embedding framework that maps each utterance…

音频与语音处理 · 电气工程与系统科学 2026-05-11 Jim O'Regan , Jens Edlund

Recently, frequency domain all-neural beamforming methods have achieved remarkable progress for multichannel speech separation. In parallel, the integration of time domain network structure and beamforming also gains significant attention.…

声音 · 计算机科学 2022-12-27 Rongzhi Gu , Shi-Xiong Zhang , Yuexian Zou , Dong Yu

In this paper, we propose a new differentiable neural network alignment mechanism for text-dependent speaker verification which uses alignment models to produce a supervector representation of an utterance. Unlike previous works with…

声音 · 计算机科学 2018-12-27 Victoria Mingote , Antonio Miguel , Alfonso Ortega , Eduardo Lleida

This paper addresses the problem of localizing audio sources using binaural measurements. We propose a supervised formulation that simultaneously localizes multiple sources at different locations. The approach is intrinsically efficient…

声音 · 计算机科学 2016-04-18 Antoine Deleforge , Radu Horaud , Yoav Schechner , Laurent Girin

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which…

音频与语音处理 · 电气工程与系统科学 2019-03-26 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversion. The contrastive…

音频与语音处理 · 电气工程与系统科学 2024-09-06 Yuying Xie , Michael Kuhlmann , Frederik Rautenberg , Zheng-Hua Tan , Reinhold Haeb-Umbach

Distant speech processing is a challenging task, especially when dealing with the cocktail party effect. Sound source separation is thus often required as a preprocessing step prior to speech recognition to improve the signal to distortion…

音频与语音处理 · 电气工程与系统科学 2020-08-06 Francois Grondin , Jean-Samuel Lauzon , Jonathan Vincent , Francois Michaud

Separation of competing speech is a key challenge in signal processing and a feat routinely performed by the human auditory brain. A long standing benchmark of the spectrogram approach to source separation is known as the ideal binary mask.…

声音 · 计算机科学 2015-03-25 Andrew J. R. Simpson

Discovering speaker independent acoustic units purely from spoken input is known to be a hard problem. In this work we propose an unsupervised speaker normalization technique prior to unit discovery. It is based on separating speaker…

音频与语音处理 · 电气工程与系统科学 2021-05-06 Thomas Glarner , Janek Ebbers , Reinhold Häb-Umbach

Although mask-based beamforming is a powerful speech enhancement approach, it often requires manual parameter tuning to handle moving speakers. Recently, this approach was augmented with an attention-based spatial covariance matrix…

音频与语音处理 · 电气工程与系统科学 2024-06-18 Marvin Tammen , Tsubasa Ochiai , Marc Delcroix , Tomohiro Nakatani , Shoko Araki , Simon Doclo

Space-time codes leverage the availability of multiple antennas to enhance the reliability of communication over wireless channels. While space-time codes have initially been designed with a focus on open-loop systems, recent technological…

信息论 · 计算机科学 2016-11-17 Che Lin , Vasanthan Raghavan , Venu Veeravalli

Sound source tracking is commonly performed using classical array-processing algorithms, while machine-learning approaches typically rely on precise source position labels that are expensive or impractical to obtain. This paper introduces a…

音频与语音处理 · 电气工程与系统科学 2026-02-12 Luan Vinícius Fiorio , Ivana Nikoloska , Bruno Defraene , Alex Young , Johan David , Ronald M. Aarts

Most neural-network based speaker-adaptive acoustic models for speech synthesis can be categorized into either layer-based or input-code approaches. Although both approaches have their own pros and cons, most existing works on speaker…

音频与语音处理 · 电气工程与系统科学 2018-10-02 Hieu-Thi Luong , Junichi Yamagishi

Incremental improvements in accuracy of Convolutional Neural Networks are usually achieved through use of deeper and more complex models trained on larger datasets. However, enlarging dataset and models increases the computation and storage…

音频与语音处理 · 电气工程与系统科学 2018-07-24 Mahdi Hajibabaei , Dengxin Dai

Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from artificially mixed spectrograms, known as Mix-and-Separate…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Tanzila Rahman , Leonid Sigal

We propose an algorithm to separate simultaneously speaking persons from each other, the "cocktail party problem", using a single microphone. Our approach involves a deep recurrent neural networks regression to a vector space that is…

声音 · 计算机科学 2017-05-22 Cory Stephenson , Patrick Callier , Abhinav Ganesh , Karl Ni

Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification), a framework that…

声音 · 计算机科学 2025-03-14 Jakaria Islam Emon , Md Abu Salek , Kazi Tamanna Alam