English
Related papers

Related papers: SDR - half-baked or well done?

200 papers

Distant speech processing is a challenging task, especially when dealing with the cocktail party effect. Sound source separation is thus often required as a preprocessing step prior to speech recognition to improve the signal to distortion…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Francois Grondin , Jean-Samuel Lauzon , Jonathan Vincent , Francois Michaud

Unsupervised denoising is a crucial challenge in real-world imaging applications. Unsupervised deep-learning methods have demonstrated impressive performance on benchmarks based on synthetic noise. However, no metrics are available to…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Adria Marcos-Morales , Matan Leibovich , Sreyas Mohan , Joshua Lawrence Vincent , Piyush Haluai , Mai Tan , Peter Crozier , Carlos Fernandez-Granda

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative models may introduce…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-19 Paul A. Bereuter , Benjamin Stahl , Mark D. Plumbley , Alois Sontacchi

This paper presents a neural method for distant speech recognition (DSR) that jointly separates and diarizes speech mixtures without supervision by isolated signals. A standard separation method for multi-talker DSR is a statistical…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Yoshiaki Bando , Tomohiko Nakamura , Shinji Watanabe

Self-supervised learning (SSL) is the latest breakthrough in speech processing, especially for label-scarce downstream tasks by leveraging massive unlabeled audio data. The noise robustness of the SSL is one of the important challenges to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-25 Hiroshi Sato , Ryo Masumura , Tsubasa Ochiai , Marc Delcroix , Takafumi Moriya , Takanori Ashihara , Kentaro Shinayama , Saki Mizuno , Mana Ihori , Tomohiro Tanaka , Nobukatsu Hojo

Sparse Representation (SR) of signals or data has a well founded theory with rigorous mathematical error bounds and proofs. SR of a signal is given by superposition of very few columns of a matrix called Dictionary, implicitly reducing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 G. Madhuri , Atul Negi

We address the problem of super-resolution of point sources from binary measurements, where random projections of the blurred measurement of the actual signal are encoded using only the sign information. The threshold used for binary…

Information Theory · Computer Science 2016-06-14 Subhadip Mukherjee , Anjany Kumar Sekuboyina , Chandra Sekhar Seelamantula

Most previously proposed dual-channel coherent-to-diffuse-ratio (CDR) estimators are based on a free-field model. When used for binaural signals, e.g., for dereverberation in binaural hearing aids, their performance may degrade due to the…

Sound · Computer Science 2015-06-12 Chengshi Zheng , Andreas Schwarz , Walter Kellermann , Xiaodong Li

A core motivation of science is to evaluate which scientific model best explains observed data. Bayesian model comparison provides a principled statistical approach to comparing scientific models and has found widespread application within…

Cosmology and Nongalactic Astrophysics · Physics 2025-06-06 Kiyam Lin , Alicja Polanska , Davide Piras , Alessio Spurio Mancini , Jason D. McEwen

Bias is a common problem inherent in recommender systems, which is entangled with users' preferences and poses a great challenge to unbiased learning. For debiasing tasks, the doubly robust (DR) method and its variants show superior…

Information Retrieval · Computer Science 2023-03-03 Haoxuan Li , Yan Lyu , Chunyuan Zheng , Peng Wu

The paper proposes an efficient, robust, and reconfigurable technique to suppress various types of noises for any sampling rate. The theoretical analyses, subjective and objective test results show that the proposed noise suppression (NS)…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-30 Jun Yang , Joshua Bingham

Channel estimation is a fundamental task in communication systems and is critical for effective demodulation. While most works deal with a simple scenario where the measurements are corrupted by the additive white Gaussian noise (AWGN),…

Signal Processing · Electrical Eng. & Systems 2024-12-10 Yifan Wang , Chengjie Yu , Jiang Zhu , Fangyong Wang , Xingbin Tu , Yan Wei , Fengzhong Qu

This study investigates the impact of integrating a dataset of disordered speech recordings ($\sim$1,000 hours) into the fine-tuning of a near state-of-the-art ASR baseline system. Contrary to what one might expect, despite the data being…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-22 Jimmy Tobin , Katrin Tomanek , Subhashini Venugopalan

Causal inference plays an important role in under standing the underlying mechanisation of the data generation process across various domains. It is challenging to estimate the average causal effect and individual causal effects from…

Data Structures and Algorithms · Computer Science 2023-01-05 Haoran Zhao , Yinghao Zhang , Debo Cheng , Chen Li , Zaiwen Feng

The conversation scenario is one of the most important and most challenging scenarios for speech processing technologies because people in conversation respond to each other in a casual style. Detecting the speech activities of each person…

Computation and Language · Computer Science 2022-08-18 Gaofeng Cheng , Yifan Chen , Runyan Yang , Qingxuan Li , Zehui Yang , Lingxuan Ye , Pengyuan Zhang , Qingqing Zhang , Lei Xie , Yanmin Qian , Kong Aik Lee , Yonghong Yan

Word Error Rate (WER) has been the predominant metric used to evaluate the performance of automatic speech recognition (ASR) systems. However, WER is sometimes not a good indicator for downstream Natural Language Understanding (NLU) tasks,…

Computation and Language · Computer Science 2021-04-07 Suyoun Kim , Abhinav Arora , Duc Le , Ching-Feng Yeh , Christian Fuegen , Ozlem Kalinli , Michael L. Seltzer

Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic counterparts for efficient model training. However, existing DD methods exhibit substantial performance degradation on long-tailed datasets. We identify…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ruixi Wu , Shaobo Wang , Jiahuan Chen , Zhiyuan Liu , Yicun Yang , Zhaorun Chen , Zekai Li , Kaixin Li , Xinming Wang , Hongzhu Yi , Kai Wang , Linfeng Zhang

The total variation filtering technique emerges as a highly effective strategy for restoring signals with discontinuities in various parts of their structure. This study presents and implements a one-dimensional signal filtering algorithm…

Optimization and Control · Mathematics 2024-10-14 Joyce Oliveira dos Santos , Francisco Márcio Barboza

Semi-supervised learning (SSL) can reduce the need for large labelled datasets by incorporating unlabelled data into the training. This is particularly interesting for semantic segmentation, where labelling data is very costly and…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Sebastian Scherer , Robin Schön , Rainer Lienhart

The cocktail party problem aims at isolating any source of interest within a complex acoustic scene, and has long inspired audio source separation research. Recent efforts have mainly focused on separating speech from noise, speech from…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-25 Darius Petermann , Gordon Wichern , Zhong-Qiu Wang , Jonathan Le Roux