English
Related papers

Related papers: iSTFTNet: Fast and Lightweight Mel-Spectrogram Voc…

200 papers

Recently, we proposed short-time Fourier transform (STFT)-based loss functions for training a neural speech waveform model. In this paper, we generalize the above framework and propose a training scheme for such models based on spectral…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-09 Shinji Takaki , Hirokazu Kameoka , Junichi Yamagishi

Neural vocoders have recently advanced waveform generation, yielding natural and expressive audio. Among these approaches, iSTFT-based vocoders have recently gained attention. They predict a complex-valued spectrogram and then synthesize…

Sound · Computer Science 2026-03-13 Hyung-Seok Oh , Deok-Hyeon Cho , Seung-Bin Kim , Seong-Whan Lee

Consumer-grade music recordings such as those captured by mobile devices typically contain distortions in the form of background noise, reverb, and microphone-induced EQ. This paper presents a deep learning approach to enhance low-quality…

Sound · Computer Science 2022-04-29 Nikhil Kandpal , Oriol Nieto , Zeyu Jin

In this study, we propose a dense frequency-time attentive network (DeFT-AN) for multichannel speech enhancement. DeFT-AN is a mask estimation network that predicts a complex spectral masking pattern for suppressing the noise and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-07 Dongheon Lee , Jung-Woo Choi

Neural network-based vocoders have recently demonstrated the powerful ability to synthesize high-quality speech. These models usually generate samples by conditioning on spectral features, such as Mel-spectrogram and fundamental frequency,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-13 Yunchao He , Yujun Wang

For training the sequence-to-sequence voice conversion model, we need to handle an issue of insufficient data about the number of speech pairs which consist of the same utterance. This study experimentally investigated the effects of…

Machine Learning · Computer Science 2020-06-16 Yeongtae Hwang , Hyemin Cho , Hongsun Yang , Dong-Ok Won , Insoo Oh , Seong-Whan Lee

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-31 Shinji Takaki , Toru Nakashika , Xin Wang , Junichi Yamagishi

Although supervised learning based on a deep neural network has recently achieved substantial improvement on speech enhancement, the existing schemes have either of two critical issues: spectrum or metric mismatches. The spectrum mismatch…

Sound · Computer Science 2020-05-12 Jaeyoung Kim , Mostafa El-Khamy , Jungwon Lee

Non-parallel voice conversion (VC) is a technique for learning mappings between source and target speeches without using a parallel corpus. Recently, cycle-consistent adversarial network (CycleGAN)-VC and CycleGAN-VC2 have shown promising…

Sound · Computer Science 2020-10-23 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Nobukatsu Hojo

In this work, we propose a new mathematical vocoder algorithm(modified spectral inversion) that generates a waveform from acoustic features without phase estimation. The main benefit of using our proposed method is that it excludes the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Hyun Gon Ryu , Jeong-Hoon Kim , Simon See

Transformer has shown promise in reinforcement learning to model time-varying features for obtaining generalized low-level robot policies on diverse robotics datasets in embodied learning. However, it still suffers from the issues of low…

Machine Learning · Computer Science 2024-12-19 Hengkai Tan , Songming Liu , Kai Ma , Chengyang Ying , Xingxing Zhang , Hang Su , Jun Zhu

With the emergence of GAN-based vocoders, the discriminator, as a crucial component, has been developed recently. In our work, we focus on improving the time-frequency based discriminator. Particularly, Short-Time Fourier Transform (STFT)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-04 Nan Xu , Zhaolong Huang , Xiao Zeng

With the development of automatic speech recognition (ASR) and text-to-speech (TTS) technology, high-quality voice conversion (VC) can be achieved by extracting source content information and target speaker information to reconstruct…

Sound · Computer Science 2023-02-24 Houjian Guo , Chaoran Liu , Carlos Toshinori Ishi , Hiroshi Ishiguro

Generative Adversarial Network (GAN) based vocoders are superior in both inference speed and synthesis quality when reconstructing an audible waveform from an acoustic representation. This study focuses on improving the discriminator for…

Sound · Computer Science 2024-04-29 Yicheng Gu , Xueyao Zhang , Liumeng Xue , Haizhou Li , Zhizheng Wu

In this paper, we address the problem of multichannel speech enhancement in the short-time Fourier transform (STFT) domain. A long short-time memory (LSTM) network takes as input a sequence of STFT coefficients associated with a frequency…

Sound · Computer Science 2020-09-24 Xiaofei LI , Radu Horaud

We propose a lightweight end-to-end text-to-speech model using multi-band generation and inverse short-time Fourier transform. Our model is based on VITS, a high-quality end-to-end text-to-speech model, but adopts two changes for more…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-22 Masaya Kawamura , Yuma Shirahata , Ryuichi Yamamoto , Kentaro Tachibana

The phase vocoder (PV) is a widely spread technique for processing audio signals. It employs a short-time Fourier transform (STFT) analysis-modify-synthesis loop and is typically used for time-scaling of signals by means of using different…

Sound · Computer Science 2022-02-16 Zdenek Prusa , Nicki Holighaus

In this paper, we propose to unify the two aspects of voice synthesis, namely text-to-speech (TTS) and vocoder, into one framework based on a pair of forward and reverse-time linear stochastic differential equations (SDE). The solutions of…

Sound · Computer Science 2022-02-01 Shoule Wu , Ziqiang Shi

Generative Adversarial Network (GAN) based vocoders are superior in inference speed and synthesis quality when reconstructing an audible waveform from an acoustic representation. This study focuses on improving the discriminator to promote…

Sound · Computer Science 2023-11-28 Yicheng Gu , Xueyao Zhang , Liumeng Xue , Zhizheng Wu

We propose Universal MelGAN, a vocoder that synthesizes high-fidelity speech in multiple domains. To preserve sound quality when the MelGAN-based structure is trained with a dataset of hundreds of speakers, we added multi-resolution…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-05 Won Jang , Dan Lim , Jaesam Yoon