English
Related papers

Related papers: UBGAN: Enhancing Coded Speech with Blind and Guide…

200 papers

UWB has a very large bandwidth in a WPAN network, which is best used for HD-video applications. Meanwhile, MBWA is a WMAN option optimized for wireless-IP in a fast moving vehicle. In this paper, we propose a practical engineering scenario…

Information Theory · Computer Science 2013-06-03 Mouhamed Abdulla , Yousef R. Shayan

In this article, we develop an end-to-end wireless communication system using deep neural networks (DNNs), in which DNNs are employed to perform several key functions, including encoding, decoding, modulation, and demodulation. However, an…

Information Theory · Computer Science 2019-03-08 Hao Ye , Le Liang , Geoffrey Ye Li , Biing-Hwang Fred Juang

Training of speech enhancement systems often does not incorporate knowledge of human perception and thus can lead to unnatural sounding results. Incorporating psychoacoustically motivated speech perception metrics as part of model training…

Sound · Computer Science 2022-06-16 George Close , Thomas Hain , Stefan Goetze

Both reverberation and additive noises degrade the speech quality and intelligibility. Weighted prediction error (WPE) method performs well on the dereverberation but with limitations. First, WPE doesn't consider the influence of the…

Sound · Computer Science 2017-08-29 Hao Li , Xueliang Zhang , Hui Zhang , Guanglai Gao

Sixth-generation (6G) wireless networks must support heterogeneous services: enhanced Mobile Broadband (eMBB) requiring 1 Tbps data rates, massive Machine-Type Communications (mMTC) supporting 10 million devices per km, and Ultra-Reliable…

Networking and Internet Architecture · Computer Science 2026-04-13 Daniel Benniah John

The advent of Large Models marks a new era in machine learning, significantly outperforming smaller models by leveraging vast datasets to capture and synthesize complex patterns. Despite these advancements, the exploration into scaling,…

Sound · Computer Science 2024-02-05 Shijia Liao , Shiyi Lan , Arun George Zachariah

Brain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-10 Jihwan Lee , Tiantian Feng , Aditya Kommineni , Sudarsana Reddy Kadiri , Shrikanth Narayanan

In underwater acoustic (UWA) communication, orthogonal frequency division multiplexing (OFDM) is commonly employed to mitigate the inter-symbol interference (ISI) caused by delay spread. However, path-specific Doppler effects in UWA…

Signal Processing · Electrical Eng. & Systems 2024-09-24 Xiaoquan You , Hengyu Zhang , Xuehan Wang , Jintao Wang

Besides the well-known classification task, these days neural networks are frequently being applied to generate or transform data, such as images and audio signals. In such tasks, the conventional loss functions like the mean squared error…

Recent advancements in deep learning led to human-level performance in single-speaker speech synthesis. However, there are still limitations in terms of speech quality when generalizing those systems into multiple-speaker models especially…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-13 Dipjyoti Paul , Yannis Pantazis , Yannis Stylianou

Noise robustness remains a critical challenge for deploying neural speech codecs in real-world acoustic scenarios where background noise is often inevitable. A key observation we make is that even slight input noise perturbations can cause…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-14 Rui-Chen Zheng , Yang Ai , Hui-Peng Du , Li-Rong Dai

Recent approaches in text-to-speech (TTS) synthesis employ neural network strategies to vocode perceptually-informed spectrogram representations directly into listenable waveforms. Such vocoding procedures create a computational bottleneck…

Sound · Computer Science 2019-07-29 Paarth Neekhara , Chris Donahue , Miller Puckette , Shlomo Dubnov , Julian McAuley

Generative speech enhancement methods based on generative adversarial networks (GANs) and diffusion models have shown promising results in various speech enhancement tasks. However, their performance in very low signal-to-noise ratio (SNR)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-29 Shrishti Saha Shetu , Emanuël A. P. Habets , Andreas Brendel

The most striking successes in image retrieval using deep hashing have mostly involved discriminative models, which require labels. In this paper, we use binary generative adversarial networks (BGAN) to embed images to binary codes in an…

Computer Vision and Pattern Recognition · Computer Science 2017-08-15 Jingkuan Song

Recently, neural networks have proven to be effective in performing speech coding task at low bitrates. However, under-utilization of intra-frame correlations and the error of quantizer specifically degrade the reconstructed audio quality.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-05 Linping Xu , Jiawei Jiang , Dejun Zhang , Xianjun Xia , Li Chen , Yijian Xiao , Piao Ding , Shenyi Song , Sixing Yin , Ferdous Sohel

Neural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on…

Sound · Computer Science 2023-12-15 George Close , William Ravenscroft , Thomas Hain , Stefan Goetze

Speech enhancement plays an essential role in improving the quality of speech signals in noisy environments. This paper investigates the efficacy of integrating Bidirectional Gated Recurrent Units (BGRU) and Transformer models for speech…

Sound · Computer Science 2025-02-26 Souliman Alghnam , Mohammad Alhussien , Khaled Shaheen

Neural speech codecs have demonstrated their ability to compress high-quality speech and audio by converting them into discrete token representations. Most existing methods utilize Residual Vector Quantization (RVQ) to encode speech into…

Sound · Computer Science 2024-10-22 Peiji Yang , Fengping Wang , Yicheng Zhong , Huawei Wei , Zhisheng Wang

In this paper, we address the design of high spectral-efficiency Barnes-Wall (BW) lattice codes which are amenable to low-complexity decoding in additive white Gaussian noise (AWGN) channels. We propose a new method of constructing complex…

Information Theory · Computer Science 2013-01-09 J. Harshan , Emanuele Viterbo , Jean-Claude Belfiore

Millimeter-wave (mmWave) radar captures are band-limited and noisy, making for difficult reconstruction of intelligible full-bandwidth speech. In this work, we propose a two-stage speech reconstruction pipeline for mmWave using a…

Sound · Computer Science 2026-02-27 Jash Karani , Adithya Chittem , Deepan Roy , Sandeep Joshi