中文
相关论文

相关论文: SDR - half-baked or well done?

200 篇论文

Many state-of-the-art neural network-based source separation systems use the averaged Signal-to-Distortion Ratio (SDR) as a training objective function. The basic SDR is, however, undefined if the network reconstructs the reference signal…

音频与语音处理 · 电气工程与系统科学 2022-04-22 Thilo von Neumann , Keisuke Kinoshita , Christoph Boeddeker , Marc Delcroix , Reinhold Haeb-Umbach

This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with…

音频与语音处理 · 电气工程与系统科学 2025-08-21 Simon Dahl Jepsen , Mads Græsbøll Christensen , Jesper Rindom Jensen

We revisit the widely used bss eval metrics for source separation with an eye out for performance. We propose a fast algorithm fixing shortcomings of publicly available implementations. First, we show that the metrics are fully specified by…

音频与语音处理 · 电气工程与系统科学 2021-10-14 Robin Scheibler

Recently, deep neural network (DNN) has made a breakthrough in monaural source enhancement. Through a training step by using a large amount of data, DNN estimates a mapping between mixed signals and clean signals. At this time, we use an…

声音 · 计算机科学 2018-06-18 Hiroaki Nakajima , Yu Takahashi , Kazunobu Kondo , Yuji Hisaminato

Usually, hearing impaired people use hearing aids which are implemented with speech enhancement algorithms. Estimation of speech and estimation of nose are the components in single channel speech enhancement system. The main objective of…

声音 · 计算机科学 2014-11-10 M. Ravichandra Kumar , B. Ravi Teja

Spatial semantic segmentation of sound scenes (S5) consists of jointly performing audio source separation and sound event classification from a multichannel audio mixture. Evaluating S5 systems with separation and classification metrics…

声音 · 计算机科学 2026-05-27 Mayank Mishra , Paul Magron , Romain Serizel

In this paper, the task of channel sounding using software defined radios (SDRs) is considered. In contrast to classical channel sounding equipment, SDRs are general purpose devices and require additional steps to be implemented when…

信号处理 · 电气工程与系统科学 2022-05-24 Julian Ahrens , Lia Ahrens , Michael Zentarra , Hans D. Schotten

Music source separation aims to extract individual sound sources (e.g., vocals, drums, guitar) from a mixed music recording. However, evaluating the quality of separated audio remains challenging, as commonly used metrics like the…

音频与语音处理 · 电气工程与系统科学 2025-10-01 Noah Jaffe , John Ashley Burgoyne

This paper introduces a practical approach for leveraging a real-time deep learning model to alternate between speech enhancement and joint speech enhancement and separation depending on whether the input mixture contains one or two active…

音频与语音处理 · 电气工程与系统科学 2023-10-17 Kashyap Patel , Anton Kovalyov , Issa Panahi

Time-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech…

In this paper we suggest a new algorithm for determination of signal-to-noise ratio (SNR). SNR is a quantitative measure widely used in science and engineering. Generally, methods for determination of SNR are based on using of…

数据分析、统计与概率 · 物理学 2016-09-30 Z. Zh. Zhanabaev , S. N. Akhtanov , E. T. Kozhagulov , B. A Karibayev

In applications that involve sensor data, a useful measure of signal-to-noise ratio (SNR) is the ratio of the root-mean-squared (RMS) signal to the RMS sensor noise. The present paper shows that, for numerical differentiation, the…

系统与控制 · 电气工程与系统科学 2025-01-28 Shashank Verma , Mohammad Almuhaihi , Dennis S. Bernstein

Deepfake audio detection has progressed rapidly with strong pre-trained encoders (e.g., WavLM, Wav2Vec2, MMS). However, performance in realistic capture conditions - background noise (domestic/office/transport), room reverberation, and…

声音 · 计算机科学 2025-12-17 Udayon Sen , Alka Luqman , Anupam Chattopadhyay

A previous signal processing algorithm that aimed to enhance spectral changes (SCE) over time showed benefit for hearing-impaired (HI) listeners to recognize speech in background noise. In this work, the previous SCE was manipulated to…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Xiang Li , Xin Tian , Henry Luo , Jinyu Qian , Xihong Wu , Dingsheng Luo , Jing Chen

In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human…

声音 · 计算机科学 2026-01-21 Lingling Dai , Andong Li , Cheng Chi , Yifan Liang , Xiaodong Li , Chengshi Zheng

It is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with a single-channel speech enhancement (SE) front-end. This is generally attributed to the processing distortions caused by the nonlinear…

音频与语音处理 · 电气工程与系统科学 2024-04-24 Tsubasa Ochiai , Kazuma Iwamoto , Marc Delcroix , Rintaro Ikeshita , Hiroshi Sato , Shoko Araki , Shigeru Katagiri

Recent neural network strategies for source separation attempt to model audio signals by processing their waveforms directly. Mean squared error (MSE) that measures the Euclidean distance between waveforms of denoised speech and the…

音频与语音处理 · 电气工程与系统科学 2018-06-05 Shrikant Venkataramani , Ryley Higa , Paris Smaragdis

Human subjective evaluation is optimal to assess speech quality for human perception. The recently introduced deep noise suppression mean opinion score (DNSMOS) metric was shown to estimate human ratings with great accuracy. The…

声音 · 计算机科学 2021-07-16 Amir Ivry , Israel Cohen , Baruch Berdugo

The distance transform (DT) and its many variations are ubiquitous tools for image processing and analysis. In many imaging scenarios, the images of interest are corrupted by noise. This has a strong negative impact on the accuracy of the…

计算机视觉与模式识别 · 计算机科学 2020-09-14 Johan Öfverstedt , Joakim Lindblad , Nataša Sladoje

In recent years time domain speech separation has excelled over frequency domain separation in single channel scenarios and noise-free environments. In this paper we dissect the gains of the time-domain audio separation network (TasNet)…

‹ 上一页 1 2 3 10 下一页 ›