English
Related papers

Related papers: ESTVocoder: An Excitation-Spectral-Transformed Neu…

200 papers

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement.…

Sound · Computer Science 2022-10-31 Shulin He , Wei Rao , Jinjiang Liu , Jun Chen , Yukai Ju , Xueliang Zhang , Yannan Wang , Shidong Shang

Speaker extraction aims to mimic humans' selective auditory attention by extracting a target speaker's voice from a multi-talker environment. It is common to perform the extraction in frequency-domain, and reconstruct the time-domain signal…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-20 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

We propose a unified approach to data-driven source-filter modeling using a single neural network for developing a neural vocoder capable of generating high-quality synthetic speech waveforms while retaining flexibility of the source-filter…

Sound · Computer Science 2021-06-29 Reo Yoneyama , Yi-Chiao Wu , Tomoki Toda

Recent advancements in neural vocoding are predominantly driven by Generative Adversarial Networks (GANs) operating in the time-domain. While effective, this approach neglects the inductive bias offered by time-frequency representations,…

Sound · Computer Science 2024-05-30 Hubert Siuzdak

Brain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-10 Jihwan Lee , Tiantian Feng , Aditya Kommineni , Sudarsana Reddy Kadiri , Shrikanth Narayanan

Recently, autoregressive neural vocoders have provided remarkable performance in generating high-fidelity speech and have been able to produce synthetic speech in real-time. However, autoregressive neural vocoders such as WaveFlow are…

Sound · Computer Science 2022-03-28 Manh Luong , Viet Anh Tran

Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation. A mel-spectrogram is extracted from the waveform by a simple,…

Target speaker extraction (TSE) aims to isolate a specific speaker's speech from a mixture using speaker enrollment as a reference. While most existing approaches are discriminative, recent generative methods for TSE achieve strong results.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Aviv Navon , Aviv Shamsian , Yael Segal-Feldman , Neta Glazer , Gil Hetz , Joseph Keshet

It is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural vocoder called InstructSing, which can converge much faster…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-11 Chang Zeng , Chunhui Wang , Xiaoxiao Miao , Jian Zhao , Zhonglin Jiang , Yong Chen

Recent advancements in speech synthesis have leveraged GAN-based networks like HiFi-GAN and BigVGAN to produce high-fidelity waveforms from mel-spectrograms. However, these networks are computationally expensive and parameter-heavy.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Yinghao Aaron Li , Cong Han , Xilin Jiang , Nima Mesgarani

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performance when speech signals are corrupted by noise. Instead of…

Computation and Language · Computer Science 2020-05-25 Yanpei Shi , Qiang Huang , Thomas Hain

This paper addresses the challenge of enhancing the realism of vocoder-generated singing voice audio by mitigating the distinguishable disparities between synthetic and real-life recordings, particularly in high-frequency spectrogram…

Sound · Computer Science 2025-08-05 Runxuan Yang , Kai Li , Guo Chen , Xiaolin Hu

Neural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features. In this study, we propose SpecGrad that adapts the diffusion noise so that…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-08 Yuma Koizumi , Heiga Zen , Kohei Yatabe , Nanxin Chen , Michiel Bacchiani

Ultra-low-bitrate speech coding is pivotal for bandwidth-constrained communication and deep compression, yet maintaining naturalness and speaker identity at such extreme bit budgets remains challenging due to pronounced information loss and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Hui-Peng Du , Yang Ai , Xiao-Hang Jiang , Yuan Tian , Zhen-Hua Ling

This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-based vocoder without requiring any additional training or…

This paper presents a novel neural vocoder named APNet which reconstructs speech waveforms from acoustic features by predicting amplitude and phase spectra directly. The APNet vocoder is composed of an amplitude spectrum predictor (ASP) and…

Sound · Computer Science 2023-05-16 Yang Ai , Zhen-Hua Ling

Current text to speech (TTS) systems usually leverage a cascaded acoustic model and vocoder pipeline with mel-spectrograms as the intermediate representations, which suffer from two limitations: 1) the acoustic model and vocoder are…

Sound · Computer Science 2022-07-12 Yanqing Liu , Ruiqing Xue , Lei He , Xu Tan , Sheng Zhao

In this paper, we present a vocoder-free framework for audio super-resolution that employs a flow matching generative model to capture the conditional distribution of complex-valued spectral coefficients. Unlike conventional two-stage…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Woongjib Choi , Sangmin Lee , Hyungseob Lim , Hong-Goo Kang

Synthesized speech from articulatory movements can have real-world use for patients with vocal cord disorders, situations requiring silent speech, or in high-noise environments. In this work, we present EMA2S, an end-to-end multimodal…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-10 Yu-Wen Chen , Kuo-Hsuan Hung , Shang-Yi Chuang , Jonathan Sherman , Wen-Chin Huang , Xugang Lu , Yu Tsao

Decoding speech from non-invasive brain signals, such as electroencephalography (EEG), has the potential to advance brain-computer interfaces (BCIs), with applications in silent communication and assistive technologies for individuals with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-30 Terrance Yu-Hao Chen , Yulin Chen , Pontus Soederhaell , Sadrishya Agrawal , Kateryna Shapovalenko