中文
相关论文

相关论文: Parallel WaveNet conditioned on VAE latent vectors

200 篇论文

We propose WaveTrainerFit, a neural vocoder that performs high-quality waveform generation from data-driven features such as SSL features. WaveTrainerFit builds upon the WaveFit vocoder, which integrates diffusion model and generative…

音频与语音处理 · 电气工程与系统科学 2026-02-06 Hien Ohnaka , Yuma Shirahata , Masaya Kawamura

Generally speaking, the main objective when training a neural speech synthesis system is to synthesize natural and expressive speech from the output layer of the neural network without much attention given to the hidden layers. However, by…

声音 · 计算机科学 2021-06-28 Hieu-Thi Luong , Junichi Yamagishi

End-to-end neural machine translation has overtaken statistical machine translation in terms of translation quality for some language pairs, specially those with large amounts of parallel data. Besides this palpable improvement, neural…

计算与语言 · 计算机科学 2017-11-16 Cristina España-Bonet , Ádám Csaba Varga , Alberto Barrón-Cedeño , Josef van Genabith

Vocoders reconstruct speech waveforms from acoustic features and play a pivotal role in modern TTS systems. Frequent-domain GAN vocoders like Vocos and APNet2 have recently seen rapid advancements, outperforming time-domain models in…

声音 · 计算机科学 2024-06-13 Yuanjun Lv , Hai Li , Ying Yan , Junhui Liu , Danming Xie , Lei Xie

Neural vocoders have recently demonstrated high quality speech synthesis, but typically require a high computational complexity. LPCNet was proposed as a way to reduce the complexity of neural synthesis by using linear prediction (LP) to…

音频与语音处理 · 电气工程与系统科学 2022-03-31 Krishna Subramani , Jean-Marc Valin , Umut Isik , Paris Smaragdis , Arvindh Krishnaswamy

Recent advances in deep learning have facilitated the design of speaker verification systems that directly input raw waveforms. For example, RawNet extracts speaker embeddings from raw waveforms, which simplifies the process pipeline and…

音频与语音处理 · 电气工程与系统科学 2020-05-08 Jee-weon Jung , Seung-bin Kim , Hye-jin Shim , Ju-ho Kim , Ha-Jin Yu

We introduce a novel sequence-to-sequence (seq2seq) voice conversion (VC) model based on the Transformer architecture with text-to-speech (TTS) pretraining. Seq2seq VC models are attractive owing to their ability to convert prosody. While…

音频与语音处理 · 电气工程与系统科学 2019-12-17 Wen-Chin Huang , Tomoki Hayashi , Yi-Chiao Wu , Hirokazu Kameoka , Tomoki Toda

In text-to-speech (TTS) and voice conversion (VC), acoustic features, such as mel spectrograms, are typically used as synthesis or conversion targets owing to their compactness and ease of learning. However, because the ultimate goal is to…

声音 · 计算机科学 2025-08-28 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Yuto Kondo

Cycle-consistent generative adversarial networks have been widely used in non-parallel voice conversion (VC). Their ability to learn mappings between source and target features without relying on parallel training data eliminates the need…

声音 · 计算机科学 2025-06-24 Dominik Wagner , Ilja Baumann , Tobias Bocklet

Recently, GAN vocoders have seen rapid progress in speech synthesis, starting to outperform autoregressive models in perceptual quality with much higher generation speed. However, autoregressive vocoders are still the common choice for…

音频与语音处理 · 电气工程与系统科学 2021-08-10 Ahmed Mustafa , Jan Büthe , Srikanth Korse , Kishan Gupta , Guillaume Fuchs , Nicola Pia

This paper presents an end-to-end text-to-speech system with low latency on a CPU, suitable for real-time applications. The system is composed of an autoregressive attention-based sequence-to-sequence acoustic model and the LPCNet vocoder…

In this paper, we propose a high-quality generative text-to-speech (TTS) system using an effective spectrum and excitation estimation method. Our previous research verified the effectiveness of the ExcitNet-based speech generation model in…

音频与语音处理 · 电气工程与系统科学 2019-05-22 Ohsung Kwon , Eunwoo Song , Jae-Min Kim , Hong-Goo Kang

Current large language models (LLMs) primarily utilize next-token prediction method for inference, which significantly impedes their processing speed. In this paper, we introduce a novel inference methodology termed next-sentence…

人工智能 · 计算机科学 2024-08-15 Hongjun An , Yifan Chen , Zhe Sun , Xuelong Li

Neural latent variable models enable the discovery of interesting structure in speech audio data. This paper presents a comparison of two different approaches which are broadly based on predicting future time-steps or auto-encoding the…

音频与语音处理 · 电气工程与系统科学 2020-10-28 Henry Zhou , Alexei Baevski , Michael Auli

We introduce Wav2Seq, the first self-supervised approach to pre-train both parts of encoder-decoder models for speech data. We induce a pseudo language as a compact discrete representation, and formulate a self-supervised pseudo speech…

计算与语言 · 计算机科学 2022-05-03 Felix Wu , Kwangyoun Kim , Shinji Watanabe , Kyu Han , Ryan McDonald , Kilian Q. Weinberger , Yoav Artzi

Many factors influence speech yielding different renditions of a given sentence. Generative models, such as variational autoencoders (VAEs), capture this variability and allow multiple renditions of the same sentence via sampling. The…

音频与语音处理 · 电气工程与系统科学 2021-06-21 Penny Karanasou , Sri Karlapati , Alexis Moinet , Arnaud Joly , Ammar Abbas , Simon Slangen , Jaime Lorenzo Trueba , Thomas Drugman

Generating sound effects that humans want is an important topic. However, there are few studies in this area for sound generation. In this study, we investigate generating sound conditioned on a text prompt and propose a novel text-to-sound…

声音 · 计算机科学 2023-05-01 Dongchao Yang , Jianwei Yu , Helin Wang , Wen Wang , Chao Weng , Yuexian Zou , Dong Yu

As parallel training data is scarce for one-shot voice conversion (VC) tasks, waveform reconstruction is typically performed by various VC systems. A typical one-shot VC system comprises a content encoder and a speaker encoder. However, two…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Songjun Cao , Qinghua Wu , Jie Chen , Jin Li , Long Ma

When the amount of parallel sentences available to train a neural machine translation is scarce, a common practice is to generate new synthetic training samples from them. A number of approaches have been proposed to produce synthetic…

This Ph.D. thesis focuses on developing a system for high-quality speech synthesis and voice conversion. Vocoder-based speech analysis, manipulation, and synthesis plays a crucial role in various kinds of statistical parametric speech…

声音 · 计算机科学 2021-01-26 Mohammed Salah Al-Radhi