English
Related papers

Related papers: A Three-class ROC for Evaluating Doubletalk Detect…

200 papers

This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio recordings. We focus on connectionist temporal classification…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-16 Takenori Yoshimura , Tomoki Hayashi , Kazuya Takeda , Shinji Watanabe

We introduce a novel method for controlling the functionality of a hands-free speech communication device which comprises a model-based acoustic echo canceller (AEC), minimum variance distortionless response (MVDR) beamformer (BF) and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-11 Thomas Haubner , Walter Kellermann

Performing an adequate evaluation of sound event detection (SED) systems is far from trivial and is still subject to ongoing research. The recently proposed polyphonic sound detection (PSD)-receiver operating characteristic (ROC) and PSD…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-01 Janek Ebbers , Romain Serizel , Reinhold Haeb-Umbach

Cross-lingual speaker verification suffers from severe language-speaker entanglement. This causes systematic degradation in the hardest scenario: correctly accepting utterances from the same speaker across different languages while…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-01 Qituan Shangguan , Junhao Du , Kunyang Peng , Feng Xue , Hui Zhang , Xinsheng Wang , Kai Yu , Shuai Wang

Neural networks have led to tremendous performance gains for single-task speech enhancement, such as noise suppression and acoustic echo cancellation (AEC). In this work, we evaluate whether it is more useful to use a single joint or…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-14 Sebastian Braun , Maria Luis Valero

Teaching with the cooperation of expert teacher and assistant teacher, which is the so-called "double-teachers classroom", i.e., the course is giving by the expert online and presented through projection screen at the classroom, and the…

Sound · Computer Science 2021-06-01 Lu Ma , Xintian Wang , Song Yang , Yaguang Gong , Zhongqin Wu

How to obtain good value estimation is one of the key problems in Reinforcement Learning (RL). Current value estimation methods, such as DDPG and TD3, suffer from unnecessary over- or underestimation bias. In this paper, we explore the…

Machine Learning · Computer Science 2021-06-08 Jiafei Lyu , Xiaoteng Ma , Jiangpeng Yan , Xiu Li

Acoustic Echo Cancellation (AEC) plays a key role in voice interaction. Due to the explicit mathematical principle and intelligent nature to accommodate conditions, adaptive filters with different types of implementations are always used…

Sound · Computer Science 2020-05-20 Lu Ma , Hua Huang , Pei Zhao , Tengrong Su

Acoustic Echo Cancellation (AEC) plays a key role in speech interaction by suppressing the echo received at microphone introduced by acoustic reverberations from loudspeakers. Since the performance of linear adaptive filter (AF) would…

Sound · Computer Science 2021-06-02 Lu Ma , Song Yang , Yaguang Gong , Zhongqin Wu

Receiver operating characteristic (ROC) and detection error tradeoff (DET) curves are two widely used evaluation metrics for speaker verification. They are equivalent since the latter can be obtained by transforming the former's true…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-22 Zhongxin Bai , Xiao-Lei Zhang , Jingdong Chen

A main challenge in applying deep learning to music processing is the availability of training data. One potential solution is Multi-task Learning, in which the model also learns to solve related auxiliary tasks on additional datasets to…

Sound · Computer Science 2018-04-06 Daniel Stoller , Sebastian Ewert , Simon Dixon

Audio-Visual Speaker Detection (AVSD) hinges on modeling both individual temporal continuity and inter-personal social context. Existing coupled architectures struggle to reconcile these tasks in shared representation spaces due to…

Multimedia · Computer Science 2026-04-17 Junhao Xiao , Shun Feng , Zhiyu Wu , Jinghan Yu , Haibiao Yao , Zhiyuan Ma , Jianjun Li , Youjun Bao , Yi Chen

The topic of deep acoustic echo control (DAEC) has seen many approaches with various model topologies in recent years. Convolutional recurrent networks (CRNs), consisting of a convolutional encoder and decoder encompassing a recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-31 Ernst Seidel , Pejman Mowlaee , Tim Fingscheidt

Voice activity detection is the task of detecting speech regions in a given audio stream or recording. First, we design a neural network combining trainable filters and recurrent layers to tackle voice activity detection directly from the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-27 Marvin Lavechin , Marie-Philippe Gill , Ruben Bousbib , Hervé Bredin , Leibny Paola Garcia-Perera

Verification bias is a well-known problem that may occur in the evaluation of predictive ability of diagnostic tests. When a binary disease status is considered, various solutions can be found in the literature to correct inference based on…

Methodology · Statistics 2023-04-10 Khanh To Duc , Monica Chiogna , Gianfranco Adimari

This paper studies dual-hop amplify-and-forward relaying system employing differential encoding and decoding over time-varying Rayleigh fading channels. First, the convectional "two-symbol" differential detection (CDD) is theoretically…

Information Theory · Computer Science 2014-07-08 M. R. Avendi , Ha H. Nguyen

Voice activity and overlapped speech detection (respectively VAD and OSD) are key pre-processing tasks for speaker diarization. The final segmentation performance highly relies on the robustness of these sub-tasks. Recent studies have shown…

Sound event detection (SED) is essential for recognizing specific sounds and their temporal locations within acoustic signals. This becomes challenging particularly for on-device applications, where computational resources are limited. To…

Sound · Computer Science 2024-02-07 Yang Xiao , Rohan Kumar Das

In this work, we propose a novel cross-talk rejection framework for a multi-channel multi-talker setup for a live multiparty interactive show. Our far-field audio setup is required to be hands-free during live interaction and comprises four…

Sound · Computer Science 2024-02-16 Hyewon Han , Naveen Kumar

Echo path delay (or ref-delay) estimation is a big challenge in acoustic echo cancellation. Different devices may introduce various ref-delay in practice. Ref-delay inconsistency slows down the convergence of adaptive filters, and also…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-12 Yi Zhang , Chengyun Deng , Shiqian Ma , Yongtao Sha , Hui Song