English
Related papers

Related papers: Sequence to Sequence Neural Speech Synthesis with …

200 papers

The rapid spread of media content synthesis technology and the potentially damaging impact of audio and video deepfakes on people's lives have raised the need to implement systems able to detect these forgeries automatically. In this work…

Sound · Computer Science 2022-11-01 Luigi Attorresi , Davide Salvi , Clara Borrelli , Paolo Bestagini , Stefano Tubaro

The prosody of a spoken utterance, including features like stress, intonation and rhythm, can significantly affect the underlying semantics, and as a consequence can also affect its textual translation. Nevertheless, prosody is rarely…

Computation and Language · Computer Science 2024-11-01 Ioannis Tsiamas , Matthias Sperber , Andrew Finch , Sarthak Garg

Automatic detection of prominence at the word and syllable-levels is critical for building computer-assisted language learning systems. It has been shown that prosody embeddings learned by the current state-of-the-art (SOTA) text-to-speech…

Computation and Language · Computer Science 2024-12-12 Anindita Mondal , Rangavajjala Sankara Bharadwaj , Jhansi Mallela , Anil Kumar Vuppala , Chiranjeevi Yarra

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g.,…

Computation and Language · Computer Science 2025-08-26 Tianxin Xie , Yan Rong , Pengfei Zhang , Wenwu Wang , Li Liu

Neural sequence-to-sequence text-to-speech synthesis (TTS), such as Tacotron-2, transforms text into high-quality speech. However, generating speech with natural prosody still remains a challenge. Yasuda et. al. show that unlike natural…

Sound · Computer Science 2021-04-12 Mahsa Elyasi , Gaurav Bharaj

The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustrate pho- netic concepts that are otherwise difficult to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-10 Frederik Rautenberg , Fritz Seebauer , Jana Wiechmann , Michael Kuhlmann , Petra Wagner , Reinhold Haeb-Umbach

Using a text description as prompt to guide the generation of text or images (e.g., GPT-3 or DALLE-2) has drawn wide attention recently. Beyond text and image generation, in this work, we explore the possibility of utilizing text…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-23 Zhifang Guo , Yichong Leng , Yihan Wu , Sheng Zhao , Xu Tan

Despite recent advances, synthetic voices often lack expressiveness due to limited prosody control in commercial text-to-speech (TTS) systems. We introduce the first end-to-end pipeline that inserts Speech Synthesis Markup Language (SSML)…

Computation and Language · Computer Science 2025-08-26 Nassima Ould Ouali , Awais Hussain Sani , Ruben Bueno , Jonah Dauvet , Tim Luka Horstmann , Eric Moulines

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-11 Ziyue Jiang , Jinglin Liu , Yi Ren , Jinzheng He , Zhenhui Ye , Shengpeng Ji , Qian Yang , Chen Zhang , Pengfei Wei , Chunfeng Wang , Xiang Yin , Zejun Ma , Zhou Zhao

We introduce a text-to-speech(TTS) framework based on a neural transducer. We use discretized semantic tokens acquired from wav2vec2.0 embeddings, which makes it easy to adopt a neural transducer for the TTS framework enjoying its monotonic…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-09 Minchan Kim , Myeonghun Jeong , Byoung Jin Choi , Dongjune Lee , Nam Soo Kim

Current state-of-the-art methods for automatic synthetic speech evaluation are based on MOS prediction neural models. Such MOS prediction models include MOSNet and LDNet that use spectral features as input, and SSL-MOS that relies on a…

The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard speech for…

Sound · Computer Science 2026-01-21 Seymanur Akti , Alexander Waibel

Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which makes it challenging when converting a paragraph of texts into…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-11 Guanghui Xu , Wei Song , Zhengchen Zhang , Chao Zhang , Xiaodong He , Bowen Zhou

While generative methods have progressed rapidly in recent years, generating expressive prosody for an utterance remains a challenging task in text-to-speech synthesis. This is particularly true for systems that model prosody explicitly…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-02 Paul Mayer , Florian Lux , Alejandro Pérez-González-de-Martos , Angelina Elizarova , Lindsey Vanderlyn , Dirk Väth , Ngoc Thang Vu

Generating speech from a face image is crucial for developing virtual humans capable of interacting using their unique voices, without relying on pre-recorded human speech. In this paper, we propose Face-StyleSpeech, a zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Minki Kang , Wooseok Han , Eunho Yang

Current synthetic speech detection (SSD) methods perform well on certain datasets but still face issues of robustness and interpretability. A possible reason is that these methods do not analyze the deficiencies of synthetic speech. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-02 Yuxiang Zhang , Zhuo Li , Jingze Lu , Wenchao Wang , Pengyuan Zhang

Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a starting point to explore pruning both spectrogram prediction…

Currently, there are increasing interests in text-to-speech (TTS) synthesis to use sequence-to-sequence models with attention. These models are end-to-end meaning that they learn both co-articulation and duration properties directly from…

Sound · Computer Science 2018-10-30 Bajibabu Bollepalli , Lauri Juvela , Paavo Alku

Text-to-Speech synthesis systems are generally evaluated using Mean Opinion Score (MOS) tests, where listeners score samples of synthetic speech on a Likert scale. A major drawback of MOS tests is that they only offer a general measure of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-07 Elijah Gutierrez , Pilar Oplustil-Gallegos , Catherine Lai

Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Devang S Ram Mohan , Vivian Hu , Tian Huey Teh , Alexandra Torresquintero , Christopher G. R. Wallis , Marlene Staib , Lorenzo Foglianti , Jiameng Gao , Simon King