中文
相关论文

相关论文: STFT spectral loss for training a neural speech wa…

200 篇论文

Speaker separation refers to isolating speech of interest in a multi-talker environment. Most methods apply real-valued Time-Frequency (T-F) masks to the mixture Short-Time Fourier Transform (STFT) to reconstruct the clean speech. Hence…

音频与语音处理 · 电气工程与系统科学 2020-04-16 Zhaoheng Ni , Michael I Mandel

Most audio processing pipelines involve transformations that act on fixed-dimensional input representations of audio. For example, when using the Short Time Fourier Transform (STFT) the DFT size specifies a fixed dimension for the input…

音频与语音处理 · 电气工程与系统科学 2022-03-28 Krishna Subramani , Paris Smaragdis

Neural waveform models such as the WaveNet are used in many recent text-to-speech systems, but the original WaveNet is quite slow in waveform generation because of its autoregressive (AR) structure. Although faster non-AR models were…

音频与语音处理 · 电气工程与系统科学 2019-04-30 Xin Wang , Shinji Takaki , Junichi Yamagishi

We propose a novel training scheme to optimize voice conversion network with a speaker identity loss function. The training scheme not only minimizes frame-level spectral loss, but also speaker identity loss. We introduce a cycle…

声音 · 计算机科学 2020-11-18 Hongqiang Du , Xiaohai Tian , Lei Xie , Haizhou Li

Many audio signal processing methods are formulated in the time-frequency (T-F) domain which is obtained by the short-time Fourier transform (STFT). The properties of the STFT are fully characterized by window function, number of frequency…

信号处理 · 电气工程与系统科学 2019-02-05 Tsubasa Kusano , Yoshiki Masuyama , Kohei Yatabe , Yasuhiro Oikawa

We propose a novel spectral generative model for image synthesis that departs radically from the common variational, adversarial, and diffusion paradigms. In our approach, images, after being flattened into one-dimensional signals, are…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Andrew Kiruluta

We present a transformer-based speech-declipping model that effectively recovers clipped signals across a wide range of input signal-to-distortion ratios (SDRs). While recent time-domain deep neural network (DNN)-based declippers have…

音频与语音处理 · 电气工程与系统科学 2024-09-20 Younghoo Kwon , Jung-Woo Choi

In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to…

Tacotron-based text-to-speech (TTS) systems directly synthesize speech from text input. Such frameworks typically consist of a feature prediction network that maps character sequences to frequency-domain acoustic features, followed by a…

音频与语音处理 · 电气工程与系统科学 2020-04-08 Rui Liu , Berrak Sisman , Feilong Bao , Guanglai Gao , Haizhou Li

The Graph Fourier Transform (GFT) has recently demonstrated promising results in speech enhancement. However, existing GFT-based speech enhancement approaches often employ fixed graph topologies to build the graph Fourier basis, whose the…

音频与语音处理 · 电气工程与系统科学 2025-10-03 Tingting Wang , Tianrui Wang , Meng Ge , Qiquan Zhang , Xi Shao

We propose a novel spectral generative modeling framework for natural language processing that jointly learns a global time varying Fourier dictionary and per token mixing coefficients, replacing the ubiquitous self attention mechanism in…

计算与语言 · 计算机科学 2025-05-02 Andrew Kiruluta

Previous speech enhancement methods focus on estimating the short-time spectrum of speech signals due to its short-term stability. However, these methods often only estimate the clean magnitude spectrum and reuse the noisy phase when…

声音 · 计算机科学 2019-10-23 Chuang Geng , Lei Wang

Variational Autoencoders (VAEs) are powerful generative models, however their generated samples are known to suffer from a characteristic blurriness, as compared to the outputs of alternative generating techniques. Extensive research…

图像与视频处理 · 电气工程与系统科学 2024-01-09 Vibhu Dalal

The problem of recovering a signal from its Fourier magnitude is of paramount importance in various fields of engineering and applied physics. Due to the absence of Fourier phase information, some form of additional information is required…

信息论 · 计算机科学 2016-05-25 Kishore Jaganathan , Yonina C. Eldar , Babak Hassibi

The nonlinear Fourier transform has the potential to overcome limits on performance and achievable data rates which arise in modern optical fiber communication systems when nonlinear interference is treated as noise. The periodic nonlinear…

信号处理 · 电气工程与系统科学 2020-12-24 Jan-Willem Goossens , Hartmut Hafermann , Yves Jaouën

Speech segmentation is an essential part of speech translation (ST) systems in real-world scenarios. Since most ST models are designed to process speech segments, long-form audio must be partitioned into shorter segments before translation.…

音频与语音处理 · 电气工程与系统科学 2024-06-18 Jaesong Lee , Soyoon Kim , Hanbyul Kim , Joon Son Chung

The Continuous Wavelet Transform (CWT) is an effective tool for feature extraction in acoustic recognition using Convolutional Neural Networks (CNNs), particularly when applied to non-stationary audio. However, its high computational cost…

音频与语音处理 · 电气工程与系统科学 2025-12-01 Dang Thoai Phan , Tuan Anh Huynh , Van Tuan Pham , Cao Minh Tran , Van Thuan Mai , Ngoc Quy Tran

Contemporary speech enhancement predominantly relies on audio transforms that are trained to reconstruct a clean speech waveform. The development of high-performing neural network sound recognition systems has raised the possibility of…

音频与语音处理 · 电气工程与系统科学 2025-11-18 Mark R. Saddler , Andrew Francl , Jenelle Feather , Kaizhi Qian , Yang Zhang , Josh H. McDermott

State-of-the-art neural network language models (NNLMs) represented by long short term memory recurrent neural networks (LSTM-RNNs) and Transformers are becoming highly complex. They are prone to overfitting and poor generalization when…

计算与语言 · 计算机科学 2022-08-30 Boyang Xue , Shoukang Hu , Junhao Xu , Mengzhe Geng , Xunying Liu , Helen Meng

We describe a sequence-to-sequence neural network which directly generates speech waveforms from text inputs. The architecture extends the Tacotron model by incorporating a normalizing flow into the autoregressive decoder loop. Output…

计算与语言 · 计算机科学 2021-02-09 Ron J. Weiss , RJ Skerry-Ryan , Eric Battenberg , Soroosh Mariooryad , Diederik P. Kingma