English
Related papers

Related papers: iSTFTNet: Fast and Lightweight Mel-Spectrogram Voc…

200 papers

The Continuous Wavelet Transform (CWT) is an effective tool for feature extraction in acoustic recognition using Convolutional Neural Networks (CNNs), particularly when applied to non-stationary audio. However, its high computational cost…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-01 Dang Thoai Phan , Tuan Anh Huynh , Van Tuan Pham , Cao Minh Tran , Van Thuan Mai , Ngoc Quy Tran

This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any vocoders for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-26 Pin-Jui Ku , Alexander H. Liu , Roman Korostik , Sung-Feng Huang , Szu-Wei Fu , Ante Jukić

We propose an optimization-based method for reconstructing a time-domain signal from a low-dimensional spectral representation such as a mel-spectrogram. Phase reconstruction has been studied to reconstruct a time-domain signal from the…

Sound · Computer Science 2023-07-25 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono

Convolutional neural networks (CNN) are widely used for speech emotion recognition (SER). In such cases, the short time fourier transform (STFT) spectrogram is the most popular choice for representing speech, which is fed as input to the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-09 Shruti Gupta , Md. Shah Fahad , Akshay Deepak

We introduce a vortex phase transform with a lenslet-array to accompany shallow, dense, ``small-brain'' neural networks for high-speed and low-light imaging. Our single-shot ptychographic approach exploits the coherent diffraction, compact…

Image and Video Processing · Electrical Eng. & Systems 2025-07-01 Baurzhan Muminov , Luat T. Vuong

Current text to speech (TTS) systems usually leverage a cascaded acoustic model and vocoder pipeline with mel-spectrograms as the intermediate representations, which suffer from two limitations: 1) the acoustic model and vocoder are…

Sound · Computer Science 2022-07-12 Yanqing Liu , Ruiqing Xue , Lei He , Xu Tan , Sheng Zhao

Mel-scale spectrum features are used in various recognition and classification tasks on speech signals. There is no reason to expect that these features are optimal for all different tasks, including speaker verification (SV). This paper…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-16 Jingyu Li , Yusheng Tian , Tan Lee

This work proposes a multichannel narrow-band speech separation network. In the short-time Fourier transform (STFT) domain, the proposed network processes each frequency independently, and all frequencies use a shared network. For each…

Sound · Computer Science 2022-12-06 Changsheng Quan , Xiaofei Li

Deep neural networks have been applied to audio spectrograms for respiratory sound classification. Existing models often treat the spectrogram as a synthetic image while overlooking its physical characteristics. In this paper, a Multi-View…

Sound · Computer Science 2024-05-31 Wentao He , Yuchen Yan , Jianfeng Ren , Ruibin Bai , Xudong Jiang

Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. We present IsoNet, a…

Sound · Computer Science 2026-05-18 Dinanath Padhya , Sajen Maharjan , Binita Adhikari , Ishwor Raj Pokharel

Most audio processing pipelines involve transformations that act on fixed-dimensional input representations of audio. For example, when using the Short Time Fourier Transform (STFT) the DFT size specifies a fixed dimension for the input…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-28 Krishna Subramani , Paris Smaragdis

Recent advances in neural network -based text-to-speech have reached human level naturalness in synthetic speech. The present sequence-to-sequence models can directly map text to mel-spectrogram acoustic features, which are convenient for…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-27 Lauri Juvela , Bajibabu Bollepalli , Junichi Yamagishi , Paavo Alku

This paper proposes an end-to-end approach for single-channel speaker-independent multi-speaker speech separation, where time-frequency (T-F) masking, the short-time Fourier transform (STFT), and its inverse are represented as layers within…

Sound · Computer Science 2018-04-30 Zhong-Qiu Wang , Jonathan Le Roux , DeLiang Wang , John R. Hershey

Short-time Fourier transform (STFT) is used as the front end of many popular successful monaural speech separation methods, such as deep clustering (DPCL), permutation invariant training (PIT) and their various variants. Since the frequency…

Sound · Computer Science 2019-02-05 Ziqiang Shi , Huibin Lin , Liu Liu , Rujie Liu , Jiqing Han

Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, the use of GAN restricts training efficiency and model…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Hui-Peng Du , Yang Ai , Rui-Chen Zheng , Ye-Xin Lu , Zhen-Hua Ling

Previously, we introduced VoiceGrad, a nonparallel voice conversion (VC) technique enabling mel-spectrogram conversion from source to target speakers using a score-based diffusion model. The concept involves training a score network to…

Sound · Computer Science 2025-09-11 Hirokazu Kameoka , Takuhiro Kaneko , Kou Tanaka , Yuto Kondo

Over the past few years, speech enhancement methods based on deep learning have greatly surpassed traditional methods based on spectral subtraction and spectral estimation. Many of these new techniques operate directly in the the short-time…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Jean-Marc Valin , Umut Isik , Neerad Phansalkar , Ritwik Giri , Karim Helwani , Arvindh Krishnaswamy

This work proposes a multichannel speech separation method with narrow-band Conformer (named NBC). The network is trained to learn to automatically exploit narrow-band speech separation information, such as spatial vector clustering of…

Sound · Computer Science 2022-07-04 Changsheng Quan , Xiaofei Li

Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT…

Sound · Computer Science 2015-09-03 Scott Wisdom , Thomas Powers , Les Atlas , James Pitton

For most of the state-of-the-art speech enhancement techniques, a spectrogram is usually preferred than the respective time-domain raw data since it reveals more compact presentation together with conspicuous temporal information over a…

Sound · Computer Science 2016-08-24 Syu-Siang Wang , Alan Chern , Yu Tsao , Jeih-weih Hung , Xugang Lu , Ying-Hui Lai , Borching Su
‹ Prev 1 3 4 5 6 7 10 Next ›