English
Related papers

Related papers: Prosodic Parameter Manipulation in TTS generated s…

200 papers

While expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual voice qualities…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-16 Frederik Rautenberg , Michael Kuhlmann , Fritz Seebauer , Jana Wiechmann , Petra Wagner , Reinhold Haeb-Umbach

The prosody of a spoken word is determined by its surrounding context. In incremental text-to-speech synthesis, where the synthesizer produces an output before it has access to the complete input, the full context is often unknown which can…

Computation and Language · Computer Science 2021-06-16 Brooke Stephenson , Thomas Hueber , Laurent Girin , Laurent Besacier

This paper proposes an effective emotional text-to-speech (TTS) system with a pre-trained language model (LM)-based emotion prediction method. Unlike conventional systems that require auxiliary inputs such as manually defined emotion…

With the development of automatic speech recognition (ASR) and text-to-speech synthesis (TTS) technique, it's intuitive to construct a voice conversion system by cascading an ASR and TTS system. In this paper, we present a ASR-TTS method…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-04 Jing-Xuan Zhang , Li-Juan Liu , Yan-Nian Chen , Ya-Jun Hu , Yuan Jiang , Zhen-Hua Ling , Li-Rong Dai

Recent TTS systems are able to generate prosodically varied and realistic speech. However, it is unclear how this prosodic variation contributes to the perception of speakers' emotional states. Here we use the recent psychological paradigm…

Human-Computer Interaction · Computer Science 2021-09-08 Pol van Rijn , Silvan Mertes , Dominik Schiller , Peter M. C. Harrison , Pauline Larrouy-Maestri , Elisabeth André , Nori Jacoby

Modern text-to-speech systems are able to produce natural and high-quality speech, but speech contains factors of variation (e.g. pitch, rhythm, loudness, timbre)\ that text alone cannot contain. In this work we move towards a speech…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-29 Giorgio Fabbro , Vladimir Golkov , Thomas Kemp , Daniel Cremers

Recently, synthesizing personalized speech by text-to-speech (TTS) application is highly demanded. But the previous TTS models require a mass of target speaker speeches for training. It is a high-cost task, and hard to record lots of…

Sound · Computer Science 2022-05-25 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Streaming TTS that receives streaming text is essential for interactive systems, yet this scheme faces two major challenges: unnatural prosody due to missing lookahead and long-form collapse due to unbounded context. We propose a…

Sound · Computer Science 2026-03-09 Changsong Liu , Tianrui Wang , Ye Ni , Yizhou Peng , Eng Siong Chng

End-to-end text-to-speech (TTS) synthesis is a method that directly converts input text to output acoustic features using a single network. A recent advance of end-to-end TTS is due to a key technique called attention mechanisms, and all…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-02 Yusuke Yasuda , Xin Wang , Junichi Yamagishi

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic method to generate…

Computer Vision and Pattern Recognition · Computer Science 2021-08-29 Xinsheng Wang , Qicong Xie , Jihua Zhu , Lei Xie , Scharenborg

In this paper, we propose a method for annotating phonemic and prosodic labels on a given audio-transcript pair, aimed at constructing Japanese text-to-speech (TTS) datasets. Our approach involves fine-tuning a large-scale pre-trained…

Computation and Language · Computer Science 2025-06-10 Rui Hu , Xiaolong Lin , Jiawang Liu , Shixi Huang , Zhenpeng Zhan

Speech synthesis and music audio generation from symbolic input differ in many aspects but share some similarities. In this study, we investigate how text-to-speech synthesis techniques can be used for piano MIDI-to-audio synthesis tasks.…

Sound · Computer Science 2022-02-25 Erica Cooper , Xin Wang , Junichi Yamagishi

Recent Text-to-Speech (TTS) systems trained on reading or acted corpora have achieved near human-level naturalness. The diversity of human speech, however, often goes beyond the coverage of these corpora. We believe the ability to handle…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-09 Li-Wei Chen , Shinji Watanabe , Alexander Rudnicky

Recent advances in Text-to-Speech (TTS) have significantly improved speech naturalness, increasing the demand for precise prosody control and mispronunciation correction. Existing approaches for prosody manipulation often depend on…

Sound · Computer Science 2025-06-03 Kyowoon Lee , Artyom Stitsyuk , Gunu Jho , Inchul Hwang , Jaesik Choi

This paper proposes an Expressive Speech Synthesis model that utilizes token-level latent prosodic variables in order to capture and control utterance-level attributes, such as character acting voice and speaking style. Current works aim to…

A Prompt-based Text-To-Speech model allows a user to control different aspects of speech, such as speaking rate and perceived gender, through natural language instruction. Although user-friendly, such approaches are on one hand constrained:…

Computation and Language · Computer Science 2025-07-14 Atli Sigurgeirsson , Simon King

While recent text to speech (TTS) models perform very well in synthesizing reading-style (e.g., audiobook) speech, it is still challenging to synthesize spontaneous-style speech (e.g., podcast or conversation), mainly because of two…

Sound · Computer Science 2021-07-07 Yuzi Yan , Xu Tan , Bohan Li , Guangyan Zhang , Tao Qin , Sheng Zhao , Yuan Shen , Wei-Qiang Zhang , Tie-Yan Liu

Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-01 Hanzhao Li , Yuke Li , Xinsheng Wang , Jingbin Hu , Qicong Xie , Shan Yang , Lei Xie

The style transfer task in Text-to-Speech refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However, most existing style transfer approaches are either based on…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-01 Wenhao Guan , Yishuang Li , Tao Li , Hukai Huang , Feng Wang , Jiayan Lin , Lingyan Huang , Lin Li , Qingyang Hong

Neural TTS has shown it can generate high quality synthesized speech. In this paper, we investigate the multi-speaker latent space to improve neural TTS for adapting the system to new speakers with only several minutes of speech or…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-04 Yan Deng , Lei He , Frank Soong
‹ Prev 1 8 9 10 Next ›