English
Related papers

Related papers: Voice Impression Control in Zero-Shot TTS

200 papers

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach for emotional TTS…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-19 Xiong Cai , Dongyang Dai , Zhiyong Wu , Xiang Li , Jingbei Li , Helen Meng

Turn-taking is a fundamental aspect of human communication where speakers convey their intention to either hold, or yield, their turn through prosodic cues. Using the recently proposed Voice Activity Projection model, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Erik Ekstedt , Siyang Wang , Éva Székely , Joakim Gustafson , Gabriel Skantze

Voice Conversion (VC) for unseen speakers, also known as zero-shot VC, is an attractive research topic as it enables a range of applications like voice customizing, animation production, and others. Recent work in this area made progress…

Sound · Computer Science 2022-06-01 Shijun Wang , Dimche Kostadinov , Damian Borth

The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often face issues such as slow inference speeds, reliance on complex…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Yinghao Aaron Li , Xilin Jiang , Cong Han , Nima Mesgarani

Although current Text-To-Speech (TTS) models are able to generate high-quality speech samples, there are still challenges in developing emotion intensity controllable TTS. Most existing TTS models achieve emotion intensity control by…

Sound · Computer Science 2024-05-28 Haoxiang Shi , Jianzong Wang , Xulong Zhang , Ning Cheng , Jun Yu , Jing Xiao

Text-to-speech (TTS) technology has achieved impressive results for widely spoken languages, yet many under-resourced languages remain challenged by limited data and linguistic complexities. In this paper, we present a novel methodology…

Sound · Computer Science 2025-04-11 Yizhong Geng , Jizhuo Xu , Zeyu Liang , Jinghan Yang , Xiaoyi Shi , Xiaoyu Shen

The rapid advancement of Zero-Shot Text-to-Speech (ZS-TTS) technology has enabled high-fidelity voice synthesis from minimal audio cues, raising significant privacy and ethical concerns. Despite the threats to voice privacy, research to…

Sound · Computer Science 2025-07-29 Taesoo Kim , Jinju Kim , Dongchan Kim , Jong Hwan Ko , Gyeong-Moon Park

In this paper, we propose reverse inference optimization (RIO), a simple and effective method designed to enhance the robustness of autoregressive-model-based zero-shot text-to-speech (TTS) systems using reinforcement learning from human…

Computation and Language · Computer Science 2024-07-03 Yuchen Hu , Chen Chen , Siyin Wang , Eng Siong Chng , Chao Zhang

The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre…

One-shot voice cloning aims to transform speaker voice and speaking style in speech synthesized from a text-to-speech (TTS) system, where only a shot recording from the target reference speech can be used. Out-of-domain transfer is still a…

Sound · Computer Science 2022-02-25 Rui Li , Dong Pu , Minnie Huang , Bill Huang

Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the target speaker in the…

Computation and Language · Computer Science 2022-07-14 Yookyung Shin , Younggun Lee , Suhee Jo , Yeongtae Hwang , Taesu Kim

Expressive text-to-speech (TTS) aims to synthesize speeches with human-like tones, moods, or even artistic attributes. Recent advancements in expressive TTS empower users with the ability to directly control synthesis style through natural…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-03 Hanglei Zhang , Yiwei Guo , Sen Liu , Xie Chen , Kai Yu

Zero-shot Text-to-Speech (TTS) voice cloning poses severe privacy risks, demanding the removal of specific speaker identities from trained TTS models. Conventional machine unlearning is insufficient in this context, as zero-shot TTS can…

Style transfer TTS has shown impressive performance in recent years. However, style control is often restricted to systems built on expressive speech recordings with discrete style categories. In practical situations, users may be…

Sound · Computer Science 2023-06-02 Guanghou Liu , Yongmao Zhang , Yi Lei , Yunlin Chen , Rui Wang , Zhifei Li , Lei Xie

Accent plays a significant role in speech communication, influencing one's capability to understand as well as conveying a person's identity. This paper introduces a novel and efficient framework for accented Text-to-Speech (TTS) synthesis…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-01 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

The availability of data in expressive styles across languages is limited, and recording sessions are costly and time consuming. To overcome these issues, we demonstrate how to build low-resource, neural text-to-speech (TTS) voices with…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-01 Giulia Comini , Goeric Huybrechts , Manuel Sam Ribeiro , Adam Gabrys , Jaime Lorenzo-Trueba

We present a new neural text to speech (TTS) method that is able to transform text to speech in voices that are sampled in the wild. Unlike other systems, our solution is able to deal with unconstrained voice samples and without requiring…

Machine Learning · Computer Science 2018-02-02 Yaniv Taigman , Lior Wolf , Adam Polyak , Eliya Nachmani

Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning communities and has broad applications in the industry. As the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-26 Xu Tan , Tao Qin , Frank Soong , Tie-Yan Liu

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-22 Ryandhimas E. Zezario , Sabato M. Siniscalchi , Hsin-Min Wang , Yu Tsao

This research paper presents a comprehensive review-based study on various Text-to-Speech (TTS) technologies. TTS technology is an important aspect of human-computer interaction, enabling machines to convert written text into audible…

Sound · Computer Science 2023-12-20 Md. Jalal Uddin Chowdhury , Ashab Hussan