English
Related papers

Related papers: Deep neural network based speech separation optimi…

200 papers

In this project, we have explored machine learning approaches for predicting hearing loss thresholds on the brain's gray matter 3D images. We have solved the problem statement in two phases. In the first phase, we used a 3D CNN model to…

Machine Learning · Computer Science 2024-05-03 Trinath Sai Subhash Reddy Pittala , Uma Maheswara R Meleti , Manasa Thatipamula

Supervised speech separation uses supervised learning algorithms to learn a mapping from an input noisy signal to an output target. With the fast development of deep learning, supervised separation has become the most important direction in…

Sound · Computer Science 2017-09-05 Shasha Xia , Hao Li , Xueliang Zhang

For voice communication, it is important to extract the speech from its noisy version without introducing unnaturally artificial noise. By studying the subband mean-squared error (MSE) of the speech for unsupervised speech enhancement…

Sound · Computer Science 2019-12-10 Andong Li , Chengshi Zheng , Xiaodong Li

The INTERSPEECH 2020 Deep Noise Suppression (DNS) Challenge is intended to promote collaborative research in real-time single-channel Speech Enhancement aimed to maximize the subjective (perceptual) quality of the enhanced speech. A typical…

Blind speech separation (BSS) aims to recover multiple speech sources from multi-channel, multi-speaker mixtures under unknown array geometry and room impulse responses. In unsupervised setup where clean target speech is not available for…

Sound · Computer Science 2025-10-13 Shulin He , Zhong-Qiu Wang

In this paper, we propose a new class of high-efficiency semantic coded transmission methods for end-to-end speech transmission over wireless channels. We name the whole system as deep speech semantic transmission (DSST). Specifically, we…

Sound · Computer Science 2022-11-07 Zixuan Xiao , Shengshi Yao , Jincheng Dai , Sixian Wang , Kai Niu , Ping Zhang

Almost half a billion people world-wide suffer from disabling hearing loss. While hearing aids can partially compensate for this, a large proportion of users struggle to understand speech in situations with background noise. Here, we…

Entropy and mutual information in neural networks provide rich information on the learning process, but they have proven difficult to compute reliably in high dimensions. Indeed, in noisy and high-dimensional data, traditional estimates in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Danqi Liao , Chen Liu , Benjamin W. Christensen , Alexander Tong , Guillaume Huguet , Guy Wolf , Maximilian Nickel , Ian Adelstein , Smita Krishnaswamy

Latent neural stochastic differential equations (SDEs) have recently emerged as a promising approach for learning generative models from stochastic time series data. However, they systematically underestimate the noise level inherent in…

Machine Learning · Computer Science 2025-06-11 Linus Heck , Maximilian Gelbrecht , Michael T. Schaub , Niklas Boers

Speech-related applications deliver inferior performance in complex noise environments. Therefore, this study primarily addresses this problem by introducing speech-enhancement (SE) systems based on deep neural networks (DNNs) applied to a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-26 Syu-Siang Wang , Yu-You Liang , Jeih-weih Hung , Yu Tsao , Hsin-Min Wang , Shih-Hau Fang

Channel estimation (CE) is one of the critical signal-processing tasks of the wireless physical layer (PHY). Recent deep learning (DL) based CE have outperformed statistical approaches such as least-square-based CE (LS) and linear minimum…

Signal Processing · Electrical Eng. & Systems 2024-03-05 Animesh Sharma , Syed Asrar Ul Haq , Sumit J. Darak

Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT…

Sound · Computer Science 2015-09-03 Scott Wisdom , Thomas Powers , Les Atlas , James Pitton

In this paper, online linear regression in environments corrupted by non-Gaussian noise (especially heavy-tailed noise) is addressed. In such environments, the error between the system output and the label also does not follow a Gaussian…

Information Theory · Computer Science 2021-05-13 Sajjad Bahrami , Ertem Tuncel

Bayesian estimation of short-time spectral amplitude is one of the most predominant approaches for the enhancement of the noise corrupted speech. The performance of these estimators are usually significantly improved when any perceptually…

Sound · Computer Science 2022-02-11 Suman Samui , Indrajit Chakrabarti , Soumya K. Ghosh

Target speech extraction (TSE) has achieved strong performance in relatively simple conditions such as one-speaker-plus-noise and two-speaker mixtures, but its performance remains unsatisfactory in noisy multi-speaker scenarios. To address…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 Ziling Huang , Junnan Wu , Lichun Fan , Zhenbo Luo , Jian Luan , Haixin Guan , Yanhua Long

Speech intelligibility assessment is essential for many speech-related applications. However, most objective intelligibility metrics are intrusive, as they require clean reference speech in addition to the degraded or processed signal for…

Sound · Computer Science 2025-12-23 Wenyu Luo , Jinhui Chen

Without the need of a clean reference, non-intrusive speech assessment methods have caught great attention for objective evaluations. Recently, deep neural network (DNN) models have been applied to build non-intrusive speech assessment…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-11 Hsin-Tien Chiang , Yi-Chiao Wu , Cheng Yu , Tomoki Toda , Hsin-Min Wang , Yih-Chun Hu , Yu Tsao

With the recent advancement in the deep learning technologies such as CNNs and GANs, there is significant improvement in the quality of the images reconstructed by deep learning based super-resolution (SR) techniques. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2018-09-05 Ram Krishna Pandey , Nabagata Saha , Samarjit Karmakar , A G Ramakrishnan

Although recent neural text-to-speech (TTS) systems have achieved high-quality speech synthesis, there are cases where a TTS system generates low-quality speech, mainly caused by limited training data or information loss during knowledge…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-26 Yeunju Choi , Youngmoon Jung , Youngjoo Suh , Hoirin Kim

In sequence prediction tasks like neural machine translation, training with cross-entropy loss often leads to models that overgeneralize and plunge into local optima. In this paper, we propose an extended loss function called \emph{dual…

Computation and Language · Computer Science 2021-04-20 Zuchao Li , Hai Zhao , Yingting Wu , Fengshun Xiao , Shu Jiang