English
Related papers

Related papers: DMOSpeech 2: Reinforcement Learning for Duration P…

200 papers

Diffusion models have achieved impressive results in generative tasks such as text-to-image synthesis, yet they often struggle to fully align outputs with nuanced user intent and maintain consistent aesthetic quality. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Dohyun Kim , Seungwoo Lyu , Seung Wook Kim , Paul Hongsuck Seo

While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from imperfect alignment. We propose…

Sound · Computer Science 2026-05-22 Jinhyeok Yang , Hyeongju Kim , Yechan Yu , Joon Byun , Frederik Bous , Juheon Lee

Tacotron-based end-to-end speech synthesis has shown remarkable voice quality. However, the rendering of prosody in the synthesized speech remains to be improved, especially for long sentences, where prosodic phrasing errors can occur…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Rui Liu , Berrak Sisman , Feilong Bao , Guanglai Gao , Haizhou Li

Modern text-to-speech synthesis pipelines typically involve multiple processing stages, each of which is designed or learnt independently from the rest. In this work, we take on the challenging task of learning to synthesise speech from…

Sound · Computer Science 2021-03-18 Jeff Donahue , Sander Dieleman , Mikołaj Bińkowski , Erich Elsen , Karen Simonyan

Flow-based generative models are widely used in text-to-speech (TTS) systems to learn the distribution of audio features (e.g., Mel-spectrograms) given the input tokens and to sample from this distribution to generate diverse utterances.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Sewade Ogun , Vincent Colotte , Emmanuel Vincent

Tokenising continuous speech into sequences of discrete tokens and modelling them with language models (LMs) has led to significant success in text-to-speech (TTS) synthesis. Although these models can generate speech with high quality and…

Sound · Computer Science 2024-08-30 Zehai Tu , Guangyan Zhang , Yiting Lu , Adaeze Adigwe , Simon King , Yiwen Guo

Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting the speech rate on speech quality, such as intelligibility…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-07 Sefik Emre Eskimez , Xiaofei Wang , Manthan Thakker , Chung-Hsien Tsai , Canrun Li , Zhen Xiao , Hemin Yang , Zirun Zhu , Min Tang , Jinyu Li , Sheng Zhao , Naoyuki Kanda

We evaluate two non-autoregressive architectures, StyleTTS2 and F5-TTS, to address the spontaneous nature of in-the-wild speech. Our models utilize flexible duration modeling to improve prosodic naturalness. To handle acoustic noise, we…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Jose Giraldo , Alex Peiró-Lilja , Rodolfo Zevallos , Cristina España-Bonet

There has been a significant progress in Text-To-Speech (TTS) synthesis technology in recent years, thanks to the advancement in neural generative modeling. However, existing methods on any-speaker adaptive TTS have achieved unsatisfactory…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-15 Minki Kang , Dongchan Min , Sung Ju Hwang

This work proposes GLM-TTS, a production-level TTS system designed for efficiency, controllability, and high-fidelity speech generation. GLM-TTS follows a two-stage architecture, consisting of a text-to-token autoregressive model and a…

With the advancement of speech synthesis technology, users have higher expectations for the naturalness and expressiveness of synthesized speech. But previous research ignores the importance of prompt selection. This study proposes a…

Sound · Computer Science 2025-04-15 Dan Luo , Chengyuan Ma , Weiqin Li , Jun Wang , Wei Chen , Zhiyong Wu

With the rapid development of deep learning techniques, the generation and counterfeiting of multimedia material are becoming increasingly straightforward to perform. At the same time, sharing fake content on the web has become so simple…

Multimedia · Computer Science 2022-09-19 Davide Salvi , Brian Hosler , Paolo Bestagini , Matthew C. Stamm , Stefano Tubaro

Recently, synthesizing personalized speech by text-to-speech (TTS) application is highly demanded. But the previous TTS models require a mass of target speaker speeches for training. It is a high-cost task, and hard to record lots of…

Sound · Computer Science 2022-05-25 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Existing zero-shot text-to-speech (TTS) systems are typically designed to process complete sentences and are constrained by the maximum duration for which they have been trained. However, in many streaming applications, texts arrive…

Sound · Computer Science 2024-10-02 Trung Dang , David Aponte , Dung Tran , Tianyi Chen , Kazuhito Koishida

Text-to-speech (TTS) has been extensively studied for generating high-quality speech with textual inputs, playing a crucial role in various real-time applications. For real-world deployment, ensuring stable and timely generation in TTS…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-23 Xiaoxue Gao , Yiming Chen , Xianghu Yue , Yu Tsao , Nancy F. Chen

Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning communities and has broad applications in the industry. As the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-26 Xu Tan , Tao Qin , Frank Soong , Tie-Yan Liu

Text-to-speech synthesis (TTS) has witnessed rapid progress in recent years, where neural methods became capable of producing audios with high naturalness. However, these efforts still suffer from two types of latencies: (a) the {\em…

Computation and Language · Computer Science 2020-10-08 Mingbo Ma , Baigong Zheng , Kaibo Liu , Renjie Zheng , Hairong Liu , Kainan Peng , Kenneth Church , Liang Huang

Diffusion models have achieved remarkable success in text-to-speech (TTS), even in zero-shot scenarios. Recent efforts aim to address the trade-off between inference speed and sound quality, often considered the primary drawback of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Changjin Han , Seokgi Lee , Gyuhyeon Nam , Gyeongsu Chae

Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are frequently used, but rarely validated…

Sound · Computer Science 2026-03-03 Christoph Minixhofer , Ondrej Klejch , Peter Bell

Recently, deep learning-based Text-to-Speech (TTS) systems have achieved high-quality speech synthesis results. Recurrent neural networks have become a standard modeling technique for sequential data in TTS systems and are widely used.…

Sound · Computer Science 2024-03-19 Ziqi Liang , Haoxiang Shi , Jiawei Wang , Keda Lu