中文
相关论文

相关论文: Quasi-Periodic Parallel WaveGAN Vocoder: A Non-aut…

200 篇论文

In this paper, we propose a non-parallel any-to-many voice conversion (VC) method termed VoiceGrad. Inspired by WaveGrad, a recently introduced novel waveform generation method, VoiceGrad is based upon the concepts of score matching and…

声音 · 计算机科学 2024-03-12 Hirokazu Kameoka , Takuhiro Kaneko , Kou Tanaka , Nobukatsu Hojo , Shogo Seki

Neural speech synthesis models have recently demonstrated the ability to synthesize high quality speech for text-to-speech and compression applications. These new models often require powerful GPUs to achieve real-time operation, so being…

音频与语音处理 · 电气工程与系统科学 2019-02-20 Jean-Marc Valin , Jan Skoglund

This paper presents a low-latency real-time (LLRT) non-parallel voice conversion (VC) framework based on cyclic variational autoencoder (CycleVAE) and multiband WaveRNN with data-driven linear prediction (MWDLP). CycleVAE is a robust…

声音 · 计算机科学 2021-07-06 Patrick Lumban Tobing , Tomoki Toda

Whispered speech lacks vocal-fold excitation, making intelligible conversion challenging. We propose WhisperVC, a three-stage framework for low-resource whisper-to-normal (W2N) conversion that decouples cross-domain alignment from speech…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Dong Liu , Juan Liu , Wei Ju , Yao Tian , Ming Li

Modern Large Language Models (LLMs) rely on Transformer self-attention, which scales quadratically with sequence length. Recent linear-time alternatives, like State Space Models (SSMs), often suffer from signal degradation over extended…

计算与语言 · 计算机科学 2026-04-07 Dejan Čugalj , Aleksandar Jevremovic

Product quantisation (PQ) is a classical method for scalable vector encoding, yet it has seen limited usage for latent representations in high-fidelity image generation. In this work, we introduce PQGAN, a quantised image autoencoder that…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Denis Zavadski , Nikita Philip Tatsch , Carsten Rother

In this work, we propose ParaNet, a non-autoregressive seq2seq model that converts text to spectrogram. It is fully convolutional and brings 46.7 times speed-up over the lightweight Deep Voice 3 at synthesis, while obtaining reasonably good…

计算与语言 · 计算机科学 2020-07-01 Kainan Peng , Wei Ping , Zhao Song , Kexin Zhao

A speech enhancement method based on probabilistic geometric approach to spectral subtraction (PGA) performed on short time magnitude spectrum is presented in this paper. A confidence parameter of noise estimation is introduced in the gain…

音频与语音处理 · 电气工程与系统科学 2018-02-15 Md Tauhidul Islam , Celia Shahnaz , Wei-Ping Zhu , M. Omair Ahmad

Generative adversarial networks (GANs) have been indicated their superiority in usage of the real-time speech synthesis. Nevertheless, most of them make use of deep convolutional layers as their backbone, which may cause the absence of…

音频与语音处理 · 电气工程与系统科学 2023-02-17 Zhenxing Lu , Mengnan He , Ruixiong Zhang , Caixia Gong

In this paper, we propose an improved LPCNet vocoder using a linear prediction (LP)-structured mixture density network (MDN). The recently proposed LPCNet vocoder has successfully achieved high-quality and lightweight speech synthesis…

音频与语音处理 · 电气工程与系统科学 2020-02-03 Min-Jae Hwang , Eunwoo Song , Ryuichi Yamamoto , Frank Soong , Hong-Goo Kang

Emotional voice conversion aims to convert the spectrum and prosody to change the emotional patterns of speech, while preserving the speaker identity and linguistic content. Many studies require parallel speech data between different…

音频与语音处理 · 电气工程与系统科学 2020-10-27 Kun Zhou , Berrak Sisman , Haizhou Li

Non-autoregressive GAN-based neural vocoders are widely used due to their fast inference speed and high perceptual quality. However, they often suffer from audible artifacts such as tonal artifacts in their generated results. Therefore, we…

音频与语音处理 · 电气工程与系统科学 2024-06-11 Hyunjae Cho , Junhyeok Lee , Wonbin Jung

In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously. The key component of this system is the Duration Informed Attention Network (DurIAN), an…

计算与语言 · 计算机科学 2019-09-09 Chengzhu Yu , Heng Lu , Na Hu , Meng Yu , Chao Weng , Kun Xu , Peng Liu , Deyi Tuo , Shiyin Kang , Guangzhi Lei , Dan Su , Dong Yu

WaveCycleGAN has recently been proposed to bridge the gap between natural and synthesized speech waveforms in statistical parametric speech synthesis and provides fast inference with a moving average model rather than an autoregressive…

声音 · 计算机科学 2019-04-10 Kou Tanaka , Hirokazu Kameoka , Takuhiro Kaneko , Nobukatsu Hojo

Audio signals are sampled at high temporal resolutions, and learning to synthesize audio requires capturing structure across a range of timescales. Generative adversarial networks (GANs) have seen wide success at generating images that are…

声音 · 计算机科学 2019-02-12 Chris Donahue , Julian McAuley , Miller Puckette

Compensation for channel mismatch and noise interference is essential for robust automatic speech recognition. Enhanced speech has been introduced into the multi-condition training of acoustic models to improve their generalization ability.…

声音 · 计算机科学 2022-11-24 Hung-Shin Lee , Pin-Yuan Chen , Yao-Fei Cheng , Yu Tsao , Hsin-Min Wang

Although voice conversion (VC) algorithms have achieved remarkable success along with the development of machine learning, superior performance is still difficult to achieve when using nonparallel data. In this paper, we propose using a…

音频与语音处理 · 电气工程与系统科学 2018-04-03 Fuming Fang , Junichi Yamagishi , Isao Echizen , Jaime Lorenzo-Trueba

Most methods of voice restoration for patients suffering from aphonia either produce whispered or monotone speech. Apart from intelligibility, this type of speech lacks expressiveness and naturalness due to the absence of pitch (whispered…

声音 · 计算机科学 2018-11-07 Santiago Pascual , Antonio Bonafonte , Joan Serrà , Jose A. Gonzalez

Emotional Voice Conversion, or emotional VC, is a technique of converting speech from one emotion state into another one, keeping the basic linguistic information and speaker identity. Previous approaches for emotional VC need parallel data…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Songxiang Liu , Yuewen Cao , Helen Meng

Recent studies have shown that text-to-speech synthesis quality can be improved by using glottal vocoding. This refers to vocoders that parameterize speech into two parts, the glottal excitation and vocal tract, that occur in the human…

音频与语音处理 · 电气工程与系统科学 2019-03-15 Bajibabu Bollepalli , Lauri Juvela , Paavo Alku