English
Related papers

Related papers: Covariance Blocking and Whitening Method for Succe…

200 papers

This paper introduces a multi-microphone method for extracting a desired speaker from a mixture involving multiple speakers and directional noise in a reverberant environment. In this work, we propose leveraging the instantaneous relative…

Sound · Computer Science 2025-02-11 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

In this paper, a multi-channel time-varying covariance matrix model for late reverberation reduction is proposed. Reflecting that variance of the late reverberation is time-varying and it depends on past speech source variance, the proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-22 Masahito Togami

Speaker diarization determines who spoke and when? in an audio stream. In this study, we propose a model-based approach for robust speaker clustering using i-vectors. The ivectors extracted from different segments of same speaker are…

Sound · Computer Science 2019-07-15 Harishchandra Dubey , Abhijeet Sangwan , John Hansen

Separation of simultaneously active multiple speakers is a difficult task in environments with strong reverberation and many background noise sources. This paper uses the relative transfer matrix (ReTM), a generalization of the relative…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-13 Wageesha N. Manamperi , Thushara D. Abhayapala

This paper addresses the problem of under-determinded speech source separation from multichannel microphone singals, i.e. the convolutive mixtures of multiple sources. The time-domain signals are first transformed to the short-time Fourier…

Sound · Computer Science 2019-04-11 Xiaofei Li , Laurent Girin , Radu Horaud

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Hongyu Wang , Hui Li , Bo Li

In conventional deep speaker embedding frameworks, the pooling layer aggregates all frame-level features over time and computes their mean and standard deviation statistics as inputs to subsequent segment-level layers. Such statistics…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-24 Rongjin Li , Weibin Zhang , Dongpeng Chen , Jintao Kang , Xiaofen Xing

An important step in speaker verification is extracting features that best characterize the speaker voice. This paper investigates a front-end processing that aims at improving the performance of speaker verification based on the SVMs…

Machine Learning · Computer Science 2013-06-13 Kawthar Yasmine Zergat , Abderrahmane Amrouche

This paper presents a Head-Related Transfer Function (HRTF)-guided framework for binaural Target Speaker Extraction (TSE) from mixtures of concurrent sources. Unlike conventional TSE methods based on Direction of Arrival (DOA) estimation or…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-18 Yoav Ellinson , Sharon Gannot

In this letter, we propose enhanced factored three way restricted Boltzmann machines (EFTW-RBMs) for speech detection. The proposed model incorporates conditional feature learning by multiplying the dynamical state of the third unit, which…

Sound · Computer Science 2017-04-24 Pengfei Sun , Jun Qin

Recently in speaker recognition, performance degradation due to the channel domain mismatched condition has been actively addressed. However, the mismatches arising from language is yet to be sufficiently addressed. This paper proposes an…

Sound · Computer Science 2017-08-29 Suwon Shon , Seongkyu Mun , Hanseok Ko

In this paper we present a new method for text-independent speaker verification that combines segmental dynamic time warping (SDTW) and the d-vector approach. The d-vectors, generated from a feed forward deep neural network trained to…

Sound · Computer Science 2018-06-27 Mohamed Adel , Mohamed Afify , Akram Gaballah

We propose an algorithm to separate simultaneously speaking persons from each other, the "cocktail party problem", using a single microphone. Our approach involves a deep recurrent neural networks regression to a vector space that is…

Sound · Computer Science 2017-05-22 Cory Stephenson , Patrick Callier , Abhinav Ganesh , Karl Ni

We address performance fairness for speaker verification using the adversarial reweighting (ARW) method. ARW is reformulated for speaker verification with metric learning, and shown to improve results across different subgroups of gender…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-09 Minho Jin , Chelsea J. -T. Ju , Zeya Chen , Yi-Chieh Liu , Jasha Droppo , Andreas Stolcke

We propose a novel speech separation model designed to separate mixtures with an unknown number of speakers. The proposed model stacks 1) a dual-path processing block that can model spectro-temporal patterns, 2) a transformer decoder-based…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-24 Younglo Lee , Shukjae Choi , Byeong-Yeol Kim , Zhong-Qiu Wang , Shinji Watanabe

The performance of speaker-related systems usually degrades heavily in practical applications largely due to the presence of background noise. To improve the robustness of such systems in unknown noisy environments, this paper proposes a…

Sound · Computer Science 2018-05-04 Siyang Song , Shuimei Zhang , Björn Schuller , Linlin Shen , Michel Valstar

This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…

Sound · Computer Science 2019-01-31 Xiaofei Li , Laurent Girin , Sharon Gannot , Radu Horaud

This paper proposes a blind estimation method based on the modulation transfer function and Schroeder model for estimating reverberation time in seven-octave bands. Therefore, the speech transmission index and five room-acoustic parameters…

Sound · Computer Science 2021-03-16 Suradej Duangpummet , Jessada Karnjana , Waree Kongprawechnon , Masashi Unoki

Multi-taper estimators provide low-variance power spectrum estimates that can be used in place of the windowed discrete Fourier transform (DFT) to extract speech features such as mel-frequency cepstral coefficients (MFCCs). Even if past…

Sound · Computer Science 2021-10-27 Xuechen Liu , Md Sahidullah , Tomi Kinnunen

This work addresses the problem of block-online processing for multi-channel speech enhancement. Such processing is vital in scenarios with moving speakers and/or when very short utterances are processed, e.g., in voice assistant scenarios.…

Sound · Computer Science 2020-05-27 Jiri Malek , Zbynek Koldovsky , Marek Bohac