English
Related papers

Related papers: Glow-WaveGAN: Learning Speech Representations from…

200 papers

Speech synthesis is used in a wide variety of industries. Nonetheless, it always sounds flat or robotic. The state of the art methods that allow for prosody control are very cumbersome to use and do not allow easy tuning. To tackle some of…

Sound · Computer Science 2021-10-08 Enrique Hortal , Rodrigo Brechard Alarcia

Text-to-Speech (TTS) models can generate natural, human-like speech across multiple languages by transforming phonemes into waveforms. However, multilingual TTS remains challenging due to discrepancies in phoneme vocabularies and variations…

Sound · Computer Science 2025-04-14 Haowei Lou , Hye-young Paik , Sheng Li , Wen Hu , Lina Yao

State-of-the-art statistical parametric speech synthesis (SPSS) generally uses a vocoder to represent speech signals and parameterize them into features for subsequent modeling. Magnitude spectrum has been a dominant feature over the years.…

Sound · Computer Science 2015-10-08 Bo Fan , Siu Wa Lee , Xiaohai Tian , Lei Xie , Minghui Dong

Unsupervised representation learning of speech has been of keen interest in recent years, which is for example evident in the wide interest of the ZeroSpeech challenges. This work presents a new method for learning frame level…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Mingjie Chen , Thomas Hain

Video-conditioned audio generation, including Video-to-Sound (V2S) and Visual Text-to-Speech (VisualTTS), has traditionally been treated as distinct tasks, leaving the potential for a unified generative framework largely underexplored. In…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-23 Xin Cheng , Yuyue Wang , Xihua Wang , Yihan Wu , Kaisi Guan , Yijing Chen , Peng Zhang , Xiaojiang Liu , Meng Cao , Ruihua Song

While large language models (LLMs) have revolutionized text-to-speech (TTS) synthesis through discrete tokenization paradigms, current architectures exhibit fundamental tensions between three critical dimensions: 1) irreversible loss of…

Computation and Language · Computer Science 2025-05-29 Yaodong Song , Hongjie Chen , Jie Lian , Yuxin Zhang , Guangmin Xia , Zehan Li , Genliang Zhao , Jian Kang , Jie Li , Yongxiang Li , Xuelong Li

Variational autoencoders (VAEs) are essential tools in end-to-end representation learning. However, the sequential text generation common pitfall with VAEs is that the model tends to ignore latent variables with a strong auto-regressive…

Machine Learning · Computer Science 2021-02-26 Yang Zhao , Ping Yu , Suchismit Mahapatra , Qinliang Su , Changyou Chen

We enhance the vanilla adversarial training method for unsupervised Automatic Speech Recognition (ASR) by a diffusion-GAN. Our model (1) injects instance noises of various intensities to the generator's output and unlabeled reference text…

Computation and Language · Computer Science 2023-03-27 Xianchao Wu

Generative adversarial network (GAN)-based neural vocoders have been widely used in audio synthesis tasks due to their high generation quality, efficient inference, and small computation footprint. However, it is still challenging to train…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Sipan Li , Songxiang Liu , Luwen Zhang , Xiang Li , Yanyao Bian , Chao Weng , Zhiyong Wu , Helen Meng

Generative deep neural networks are widely used for speech synthesis, but most existing models directly generate waveforms or spectral outputs. Humans, however, produce speech by controlling articulators, which results in the production of…

Sound · Computer Science 2023-05-10 Gašper Beguš , Alan Zhou , Peter Wu , Gopala K Anumanchipalli

While pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and…

Sound · Computer Science 2025-01-09 Chien-Chun Wang , Li-Wei Chen , Cheng-Kang Chou , Hung-Shin Lee , Berlin Chen , Hsin-Min Wang

Many deep generative models are defined as a push-forward of a Gaussian measure by a continuous generator, such as Generative Adversarial Networks (GANs) or Variational Auto-Encoders (VAEs). This work explores the latent space of such deep…

Machine Learning · Computer Science 2023-05-16 Thibaut Issenhuth , Ugo Tanielian , Jérémie Mary , David Picard

The generative adversarial network (GAN) has shown its outstanding capability in improving Non-Autoregressive TTS (NAR-TTS) by adversarially training it with an extra model that discriminates between the real and the generated speech. To…

Sound · Computer Science 2022-03-23 Haohan Guo , Hui Lu , Xixin Wu , Helen Meng

Due to strict rate and reliability demands, wireless image transmission remains difficult for both classical layered designs and joint source-channel coding (JSCC), especially under low latency. Diffusion-based generative decoders can…

Machine Learning · Computer Science 2026-01-13 Jingwen Fu , Ming Xiao , Mikael Skoglund , Dong In Kim

Video-to-speech is the process of reconstructing the audio speech from a video of a spoken utterance. Previous approaches to this task have relied on a two-step process where an intermediate representation is inferred from the video, and is…

Machine Learning · Computer Science 2022-08-17 Rodrigo Mira , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Björn W. Schuller , Maja Pantic

In recent years, speech emotion recognition (SER) has been used in wide ranging applications, from healthcare to the commercial sector. In addition to signal processing approaches, methods for SER now also use deep learning techniques which…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-29 Sneha Das , Nicole Nadine Lønfeldt , Anne Katrine Pagsberg , Line H. Clemmensen

The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on generative flow matching, our prior work on generative STFT…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Simon Welker , Tal Peer , Timo Gerkmann

Talking head synthesis has emerged as a prominent research topic in computer graphics and multimedia, yet most existing methods often struggle to strike a balance between generation quality and computational efficiency, particularly under…

Graphics · Computer Science 2025-06-30 Shuai Shen , Wanhua Li , Yunpeng Zhang , Yap-Peng Tan , Jiwen Lu

Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models aimed to address these limitations and improve fidelity. However, they still face…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Seyeon Kim , Siyoon Jin , Jihye Park , Kihong Kim , Jiyoung Kim , Jisu Nam , Seungryong Kim

In this paper, we present a vocoder-free framework for audio super-resolution that employs a flow matching generative model to capture the conditional distribution of complex-valued spectral coefficients. Unlike conventional two-stage…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Woongjib Choi , Sangmin Lee , Hyungseob Lim , Hong-Goo Kang
‹ Prev 1 8 9 10 Next ›