English
Related papers

Related papers: QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on…

200 papers

Deriving a good model for multitalker babble noise can facilitate different speech processing algorithms, e.g. noise reduction, to reduce the so-called cocktail party difficulty. In the available systems, the fact that the babble waveform…

Sound · Computer Science 2017-09-19 Nasser Mohammadiha , Arne Leijon

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier transform(STFT)…

Sound · Computer Science 2019-01-03 Xingjian Du , Mengyao Zhu , Xuan Shi , Xinpeng Zhang , Wen Zhang , Jingdong Chen

The Transformer-based Whisper model has achieved state-of-the-art performance in Automatic Speech Recognition (ASR). However, its Multi-Head Attention (MHA) mechanism results in significant GPU memory consumption due to the linearly growing…

Sound · Computer Science 2026-03-03 Sen Zhang , Jianguo Wei , Wenhuan Lu , Xianghu Yue , Wei Li , Qiang Li , Pengcheng Zhao , Ming Cai , Luo Si

We propose an information theoretic framework for quantitative assessment of acoustic modeling for hidden Markov model (HMM) based automatic speech recognition (ASR). Acoustic modeling yields the probabilities of HMM sub-word states for a…

Sound · Computer Science 2017-11-09 Pranay Dighe , Afsaneh Asaei , Hervé Bourlard

Neural network applications generally benefit from larger-sized models, but for current speech enhancement models, larger scale networks often suffer from decreased robustness to the variety of real-world use cases beyond what is…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-12 Umut Isik , Ritwik Giri , Neerad Phansalkar , Jean-Marc Valin , Karim Helwani , Arvindh Krishnaswamy

This work focuses on full-body co-speech gesture generation. Existing methods typically employ an autoregressive model accompanied by vector-quantized tokens for gesture generation, which results in information loss and compromises the…

Graphics · Computer Science 2025-03-19 Binjie Liu , Lina Liu , Sanyi Zhang , Songen Gu , Yihao Zhi , Tianyi Zhu , Lei Yang , Long Ye

Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures with skip connections,…

Sound · Computer Science 2024-12-02 Shengkui Zhao , Trung Hieu Nguyen , Bin Ma

Speech Language Models (SpeechLMs) model tokenized speech to capture both semantic and acoustic information. When neural audio codecs based on Residual Vector Quantization (RVQ) are used as audio tokenizers, they produce multiple discrete…

Computation and Language · Computer Science 2026-03-06 Issa Sugiura , Shuhei Kurita , Yusuke Oda , Ryuichiro Higashinaka

Most GAN(Generative Adversarial Network)-based approaches towards high-fidelity waveform generation heavily rely on discriminators to improve their performance. However, GAN methods introduce much uncertainty into the generation process and…

Sound · Computer Science 2022-03-22 Shengyuan Xu , Wenxiao Zhao , Jing Guo

Large-scale mobile communication systems tend to contain legacy transmission channels with narrowband bottlenecks, resulting in characteristic "telephone-quality" audio. While higher quality codecs exist, due to the scale and heterogeneity…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-12 Archit Gupta , Brendan Shillingford , Yannis Assael , Thomas C. Walters

Neural audio codecs, neural networks which compress a waveform into discrete tokens, play a crucial role in the recent development of audio generative models. State-of-the-art codecs rely on the end-to-end training of an autoencoder and a…

Sound · Computer Science 2025-03-26 Zineb Lahrichi , Gaëtan Hadjeres , Gael Richard , Geoffroy Peeters

Neural Audio Codecs, initially designed as a compression technique, have gained more attention recently for speech generation. Codec models represent each audio frame as a sequence of tokens, i.e., discrete embeddings. The discrete and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-31 Alexander H. Liu , Qirui Wang , Yuan Gong , James Glass

This paper introduces a novel framework integrating nonlinear acoustic computing and reinforcement learning to enhance advanced human-robot interaction under complex noise and reverberation. Leveraging physically informed wave equations…

Robotics · Computer Science 2025-05-07 Xiaoliang Chen , Xin Yu , Le Chang , Yunhe Huang , Jiashuai He , Shibo Zhang , Jin Li , Likai Lin , Ziyu Zeng , Xianling Tu , Shuyu Zhang

Decoding continuous speech from intracortical recordings is a central challenge for brain-computer interfaces (BCIs), with transformative potential for individuals with conditions that impair their ability to speak. While recent…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-17 Tommaso Boccato , Michal Olak , Matteo Ferrante

Speech-based depression detection has shown promise as an objective diagnostic tool, yet the cross-linguistic robustness of acoustic markers and their neurobiological underpinnings remain underexplored. This study extends Cross-Data…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-07 Fuxiang Tao , Dongwei Li , Shuning Tang , Xuri Ge , Wei Ma , Anna Esposito , Alessandro Vinciarelli

Diffusion models have recently been shown to be relevant for high-quality speech generation. Most work has been focused on generating spectrograms, and as such, they further require a subsequent model to convert the spectrogram to a…

Sound · Computer Science 2024-03-12 Roi Benita , Michael Elad , Joseph Keshet

To date, various speech technology systems have adopted the vocoder approach, a method for synthesizing speech waveform that shows a major role in the performance of statistical parametric speech synthesis. WaveNet one of the best models…

Sound · Computer Science 2021-06-15 Mohammed Salah Al-Radhi , Tamás Gábor Csapó , Csaba Zainkó , Géza Németh

The prosodic aspects of speech signals produced by current text-to-speech systems are typically averaged over training material, and as such lack the variety and liveliness found in natural speech. To avoid monotony and averaged prosody…

Computation and Language · Computer Science 2019-06-05 Vincent Wan , Chun-an Chan , Tom Kenter , Jakub Vit , Rob Clark

Dysarthria speech contains the pathological characteristics of vocal tract and vocal fold, but so far, they have not yet been included in traditional acoustic feature sets. Moreover, the nonlinearity and non-stationarity of speech have been…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Ting Zhu , Shufei Duan , Camille Dingam , Huizhi Liang , Wei Zhang

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information loss. Even with…

‹ Prev 1 4 5 6 7 8 10 Next ›