English
Related papers

Related papers: Synthesizing Personalized Non-speech Vocalization …

200 papers

Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-syncing, being…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Georgios Milis , Panagiotis P. Filntisis , Anastasios Roussos , Petros Maragos

Deep learning models are becoming predominant in many fields of machine learning. Text-to-Speech (TTS), the process of synthesizing artificial speech from text, is no exception. To this end, a deep neural network is usually trained using a…

Sound · Computer Science 2021-02-11 Giuseppe Ruggiero , Enrico Zovato , Luigi Di Caro , Vincent Pollet

In recent years, the remarkable advancements in deep neural networks have brought tremendous convenience. However, the training process of a highly effective model necessitates a substantial quantity of samples, which brings huge potential…

Sound · Computer Science 2024-09-13 Zhisheng Zhang , Pengyang Huang

By representing speaker characteristic as a single fixed-length vector extracted solely from speech, we can train a neural multi-speaker speech synthesis model by conditioning the model on those vectors. This model can also be adapted to…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-09 Hieu-Thi Luong , Junichi Yamagishi

Speech Language Models (SLMs) have recently emerged as a unified paradigm for addressing a wide range of speech-related tasks, including text-to-speech (TTS), speech enhancement (SE), and automatic speech recognition (ASR). However, the…

Sound · Computer Science 2025-12-17 Yiwen Zhao , Jiatong Shi , Jinchuan Tian , Yuxun Tang , Jiarui Hai , Jionghao Han , Shinji Watanabe

We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled representations that…

We propose UnitSpeech, a speaker-adaptive speech synthesis method that fine-tunes a diffusion-based text-to-speech (TTS) model using minimal untranscribed data. To achieve this, we use the self-supervised unit representation as a pseudo…

Sound · Computer Science 2023-06-29 Heeseung Kim , Sungwon Kim , Jiheum Yeom , Sungroh Yoon

Self-supervised speech representation learning has become essential for extracting meaningful features from untranscribed audio. Recent advances highlight the potential of deriving discrete symbols from the features correlated with…

Computation and Language · Computer Science 2024-09-17 Ryota Komatsu , Takahiro Shinozaki

Most existing neural-based text-to-speech methods rely on extensive datasets and face challenges under low-resource condition. In this paper, we introduce a novel semi-supervised text-to-speech synthesis model that learns from both paired…

Sound · Computer Science 2024-02-05 Jianzong Wang , Pengcheng Li , Xulong Zhang , Ning Cheng , Jing Xiao

Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, these approaches lack…

Sound · Computer Science 2021-02-02 Paarth Neekhara , Shehzeen Hussain , Shlomo Dubnov , Farinaz Koushanfar , Julian McAuley

The goal of this contribution is to use a parametric speech synthesis system for reducing background noise and other interferences from recorded speech signals. In a first step, Hidden Markov Models of the synthesis system are trained. Two…

Sound · Computer Science 2017-07-06 Daniel Dzibela , Armin Sehr

Speech representation learning has improved both speech understanding and speech synthesis tasks for single language. However, its ability in cross-lingual scenarios has not been explored. In this paper, we extend the pretraining method for…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-06 Xiaoran Fan , Chao Pang , Tian Yuan , He Bai , Renjie Zheng , Pengfei Zhu , Shuohuan Wang , Junkun Chen , Zeyu Chen , Liang Huang , Yu Sun , Hua Wu

Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-18 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

The potential of synthetic data in text-to-speech (TTS) model training has gained increasing attention, yet its rationality and effectiveness require systematic validation. In this study, we systematically investigate the feasibility of…

Sound · Computer Science 2025-12-22 Tingxiao Zhou , Leying Zhang , Zhengyang Chen , Yanmin Qian

Self-supervised learning (SSL) speech models, which can serve as powerful upstream models to extract meaningful speech representations, have achieved unprecedented success in speech representation learning. However, their effectiveness on…

Sound · Computer Science 2023-02-01 Tung-Yu Wu , Chen-An Li , Tzu-Han Lin , Tsu-Yuan Hsu , Hung-Yi Lee

This paper proposes a new architecture for speaker adaptation of multi-speaker neural-network speech synthesis systems, in which an unseen speaker's voice can be built using a relatively small amount of speech data without transcriptions.…

Audio and Speech Processing · Electrical Eng. & Systems 2018-08-21 Hieu-Thi Luong , Junichi Yamagishi

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

Synthesizing the voices of unseen speakers remains a persisting challenge in multi-speaker text-to-speech (TTS). Existing methods model speaker characteristics through speaker conditioning during training, leading to increased model…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-18 Ismail Rasim Ulgen , Shreeram Suresh Chandra , Junchen Lu , Berrak Sisman

In this work, we propose a speaker anonymization pipeline that leverages high quality automatic speech recognition and synthesis systems to generate speech conditioned on phonetic transcriptions and anonymized speaker embeddings. Using…

Sound · Computer Science 2022-07-12 Sarina Meyer , Florian Lux , Pavel Denisov , Julia Koch , Pascal Tilli , Ngoc Thang Vu

Non-Verbal Vocalisations (NVVs) are short `non-word' utterances without proper linguistic (semantic) meaning but conveying connotations -- be this emotions/affects or other paralinguistic information. We start this contribution with a…

Sound · Computer Science 2025-08-05 Anton Batliner , Shahin Amiriparian , Björn W. Schuller