English
Related papers

Related papers: iSTFTNet: Fast and Lightweight Mel-Spectrogram Voc…

200 papers

Recent advancements in speech synthesis have leveraged GAN-based networks like HiFi-GAN and BigVGAN to produce high-fidelity waveforms from mel-spectrograms. However, these networks are computationally expensive and parameter-heavy.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-19 Yinghao Aaron Li , Cong Han , Xilin Jiang , Nima Mesgarani

The inverse short-time Fourier transform network (iSTFTNet) has garnered attention owing to its fast, lightweight, and high-fidelity speech synthesis. It obtains these characteristics using a fast and lightweight 1D CNN as the backbone and…

Sound · Computer Science 2023-08-15 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Shogo Seki

Signal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propose a mel-spectrogram inversion method based on a rigorous…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono

Online multichannel speech enhancement has been intensively studied recently. Though Mel-scale frequency is more matched with human auditory perception and computationally efficient than linear frequency, few works are implemented in a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 Yujie Yang , Bing Yang , Xiaofei Li

This paper presents FastFit, a novel neural vocoder architecture that replaces the U-Net encoder with multiple short-time Fourier transforms (STFTs) to achieve faster generation rates without sacrificing sample quality. We replaced each…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Won Jang , Dan Lim , Heayoung Park

Most state-of-the-art Text-to-Speech systems use the mel-spectrogram as an intermediate representation, to decompose the task into acoustic modelling and waveform generation. A mel-spectrogram is extracted from the waveform by a simple,…

The short-time Fourier transform (STFT) usually computes the same number of frequency components as the frame length while overlapping adjacent time frames by more than half. As a result, the number of components of a spectrogram matrix…

Signal Processing · Electrical Eng. & Systems 2020-10-29 Daichi Kitahara

This paper proposes a novel neural denoising vocoder that can generate clean speech waveforms from noisy mel-spectrograms. The proposed neural denoising vocoder consists of two components, i.e., a spectrum predictor and a enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-20 Hui-Peng Du , Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

Vocoders are models capable of transforming a low-dimensional spectral representation of an audio signal, typically the mel spectrogram, to a waveform. Modern speech generation pipelines use a vocoder as their final component. Recent…

Sound · Computer Science 2022-08-29 Bruno Di Giorgi , Mark Levy , Richard Sharp

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder successfully generates waveform sequences using a fast…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-20 Eunwoo Song , Ryuichi Yamamoto , Min-Jae Hwang , Jin-Seob Kim , Ohsung Kwon , Jae-Min Kim

Most neural vocoders employ band-limited mel-spectrograms to generate waveforms. If full-band spectral features are used as the input, the vocoder can be provided with as much acoustic information as possible. However, in some models…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Won Jang , Dan Lim , Jaesam Yoon , Bongwan Kim , Juntae Kim

Flow matching offers a robust and stable approach to training diffusion models. However, directly applying flow matching to neural vocoders can result in subpar audio quality. In this work, we present WaveFM, a reparameterized flow matching…

Sound · Computer Science 2025-03-24 Tianze Luo , Xingchen Miao , Wenbo Duan

This paper proposes a novel bidirectional neural vocoder, named BiVocoder, capable both of feature extraction and reverse waveform generation within the short-time Fourier transform (STFT) domain. For feature extraction, the BiVocoder takes…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-05 Hui-Peng Du , Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

Acoustic features play an important role in improving the quality of the synthesised speech. Currently, the Mel spectrogram is a widely employed acoustic feature in most acoustic models. However, due to the fine-grained loss caused by its…

Sound · Computer Science 2024-07-11 Guoqiang Hu , Huaning Tan , Ruilai Li

In speech synthesis and speech enhancement systems, melspectrograms need to be precise in acoustic representations. However, the generated spectrograms are over-smooth, that could not produce high quality synthesized speech. Inspired by…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-04 Leyuan Sheng , Dong-Yan Huang , Evgeniy N. Pavlovskiy

Neural network based end-to-end text to speech (TTS) has significantly improved the quality of synthesized speech. Prominent methods (e.g., Tacotron 2) usually first generate mel-spectrogram from text, and then synthesize speech from the…

Computation and Language · Computer Science 2019-11-21 Yi Ren , Yangjun Ruan , Xu Tan , Tao Qin , Sheng Zhao , Zhou Zhao , Tie-Yan Liu

Precise control over speech characteristics, such as pitch, duration, and speech rate, remains a significant challenge in the field of voice conversion. The ability to manipulate parameters like pitch and syllable rate is an important…

Sound · Computer Science 2025-07-08 Mathilde Abrassart , Nicolas Obin , Axel Roebel

This paper proposes ESTVocoder, a novel excitation-spectral-transformed neural vocoder within the framework of source-filter theory. The ESTVocoder transforms the amplitude and phase spectra of the excitation into the corresponding speech…

Sound · Computer Science 2024-11-19 Xiao-Hang Jiang , Hui-Peng Du , Yang Ai , Ye-Xin Lu , Zhen-Hua Ling

The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on generative flow matching, our prior work on generative STFT…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Simon Welker , Tal Peer , Timo Gerkmann

This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-based vocoder without requiring any additional training or…

‹ Prev 1 2 3 10 Next ›