English
Related papers

Related papers: Towards Expressive Zero-Shot Speech Synthesis with…

200 papers

This paper presents an expressive speech synthesis architecture for modeling and controlling the speaking style at a word level. It attempts to learn word-level stylistic and prosodic representations of the speech data, with the aid of two…

Sound · Computer Science 2021-11-22 Konstantinos Klapsas , Nikolaos Ellinas , June Sig Sung , Hyoungmin Park , Spyros Raptis

Voice imitation aims to transform source speech to match a reference speaker's timbre and speaking style while preserving linguistic content. A straightforward approach is to train on triplets of (source, reference, target), where source…

Sound · Computer Science 2026-04-21 Tao Feng , Yuxiang Wang , Yuancheng Wang , Xueyao Zhang , Dekun Chen , Chaoren Wang , Xun Guan , Zhizheng Wu

Electromyography-to-Speech (ETS) conversion has demonstrated its potential for silent speech interfaces by generating audible speech from Electromyography (EMG) signals during silent articulations. ETS models usually consist of an EMG…

Sound · Computer Science 2024-05-15 Zhao Ren , Kevin Scheck , Qinhan Hou , Stefano van Gogh , Michael Wand , Tanja Schultz

Although voice conversion (VC) systems have shown a remarkable ability to transfer voice style, existing methods still have an inaccurate pitch and low speaker adaptation quality. To address these challenges, we introduce Diff-HierVC, a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-09 Ha-Yeong Choi , Sang-Hoon Lee , Seong-Whan Lee

Speech super-resolution (SR) is the task that restores high-resolution speech from low-resolution input. Existing models employ simulated data and constrained experimental settings, which limit generalization to real-world SR. Predictive…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-26 Heming Wang , Eric W. Healy , DeLiang Wang

Zero-shot voice conversion aims to transfer the voice of a source speaker to that of a speaker unseen during training, while preserving the content information. Although various methods have been proposed to reconstruct speaker information…

Sound · Computer Science 2024-08-22 Anastasia Avdeeva , Aleksei Gusev

Modern neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the prosody of generated utterances often represents the average prosodic style of the database instead of having wide…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-16 Tuomo Raitio , Ramya Rasipuram , Dan Castellani

Prominences and boundaries are the essential constituents of prosodic structure in speech. They provide for means to chunk the speech stream into linguistically relevant units by providing them with relative saliences and demarcating them…

Computation and Language · Computer Science 2015-10-08 Antti Suni , Daniel Aalto , Martti Vainio

We propose a method for learning de-identified prosody representations from raw audio using a contrastive self-supervised signal. Whereas prior work has relied on conditioning models on bottlenecks, we introduce a set of inductive biases…

Computation and Language · Computer Science 2021-07-20 Jack Weston , Raphael Lenain , Udeepa Meepegama , Emil Fristed

Scaling text-to-speech to a large and wild dataset has been proven to be highly effective in achieving timbre and speech style generalization, particularly in zero-shot TTS. However, previous works usually encode speech into latent using…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-07 Ziyue Jiang , Yi Ren , Zhenhui Ye , Jinglin Liu , Chen Zhang , Qian Yang , Shengpeng Ji , Rongjie Huang , Chunfeng Wang , Xiang Yin , Zejun Ma , Zhou Zhao

Controllable speech synthesis refers to the precise control of speaking style by manipulating specific prosodic and paralinguistic attributes, such as gender, volume, speech rate, pitch, and pitch fluctuation. With the integration of…

Artificial Intelligence · Computer Science 2025-10-01 Ziyu Zhang , Hanzhao Li , Jingbin Hu , Wenhao Li , Lei Xie

Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This underexplored…

Despite remarkable advancements in recent voice conversion (VC) systems, enhancing speaker similarity in zero-shot scenarios remains challenging. This challenge arises from the difficulty of generalizing and adapting speaker characteristics…

Sound · Computer Science 2025-01-30 Ha-Yeong Choi , Jaehan Park

The style of the speech varies from person to person and every person exhibits his or her own style of speaking that is determined by the language, geography, culture and other factors. Style is best captured by prosody of a signal. High…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-15 Neeraj Kumar , Srishti Goel , Ankur Narang , Brejesh Lall

We introduce DISSC, a novel, lightweight method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. Unlike DISSC, most voice conversion (VC) methods focus primarily on timbre, and…

Sound · Computer Science 2023-10-20 Gallil Maimon , Yossi Adi

Expressive speech synthesis requires vibrant prosody and well-timed pauses. We propose an effective strategy to augment a small dataset to train an expressive end-to-end Text-to-Speech model. We merge audios of emotionally congruent text…

Sound · Computer Science 2026-02-12 Raymond Chung

We present SALAD, a zero-shot TTS autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuous representations for the next time step. We compare our…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-25 Arnon Turetzky , Avihu Dekel , Nimrod Shabtay , Slava Shechtman , David Haws , Hagai Aronowitz , Ron Hoory , Yossi Adi

Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech…

Sound · Computer Science 2023-12-19 Chunyu Qiang , Hao Li , Yixin Tian , Yi Zhao , Ying Zhang , Longbiao Wang , Jianwu Dang

Modeling the rich prosodic variations inherent in human speech is essential for generating natural-sounding speech. While speaker embeddings are commonly used as conditioning inputs in personalized speech generation, they are typically…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-22 Ismail Rasim Ulgen , John H. L. Hansen , Carlos Busso , Berrak Sisman