English
Related papers

Related papers: Latent linguistic embedding for cross-lingual text…

200 papers

When comparing speech sounds across languages, scholars often make use of feature representations of individual sounds in order to determine fine-grained sound similarities. Although binary feature systems for large numbers of speech sounds…

Computation and Language · Computer Science 2024-05-08 Arne Rubehn , Jessica Nieder , Robert Forkel , Johann-Mattis List

Neural text-to-speech (TTS) has achieved human-like synthetic speech for single-speaker, single-language synthesis. Multilingual TTS systems are limited to resource-rich languages due to the lack of large paired text and studio-quality…

Although high-fidelity speech can be obtained for intralingual speech synthesis, cross-lingual text-to-speech (CTTS) is still far from satisfactory as it is difficult to accurately retain the speaker timbres(i.e. speaker similarity) and…

Sound · Computer Science 2023-06-27 Sen Liu , Yiwei Guo , Chenpeng Du , Xie Chen , Kai Yu

In this paper we investigate cross-lingual Text-To-Speech (TTS) synthesis through the lens of adapters, in the context of lightweight TTS systems. In particular, we compare the tasks of unseen speaker and language adaptation with the goal…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-26 Alessio Falai , Ziyao Zhang , Akos Gangoly

Creating realistic and natural-sounding synthetic speech remains a big challenge for voice identities unseen during training. As there is growing interest in synthesizing voices of new speakers, here we investigate the ability of…

In this work, we propose a joint system combining a talking face generation system with a text-to-speech system that can generate multilingual talking face videos from only the text input. Our system can synthesize natural multilingual…

Computer Vision and Pattern Recognition · Computer Science 2022-05-16 Hyoung-Kyu Song , Sang Hoon Woo , Junhyeok Lee , Seungmin Yang , Hyunjae Cho , Youseong Lee , Dongho Choi , Kang-wook Kim

We present a meta-learning approach for adaptive text-to-speech (TTS) with few data. During training, we learn a multi-speaker model using a shared conditional WaveNet core and independent learned embeddings for each speaker. The aim of…

Text-to-speech (TTS) acoustic models map linguistic features into an acoustic representation out of which an audible waveform is generated. The latest and most natural TTS systems build a direct mapping between linguistic and waveform…

Sound · Computer Science 2019-09-24 David Álvarez , Santiago Pascual , Antonio Bonafonte

In this work, we explore multiple architectures and training procedures for developing a multi-speaker and multi-lingual neural TTS system with the goals of a) improving the quality when the available data in the target language is limited…

Computation and Language · Computer Science 2021-08-18 Javier Latorre , Charlotte Bailleul , Tuuli Morrill , Alistair Conkie , Yannis Stylianou

Modern voice cloning, also known as zero-shot text-to-speech (TTS), can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing.…

Sound · Computer Science 2026-05-26 Ruinan Jin , Xinting Liao , Hanlin Yu , Deval Pandya , Xiaoxiao Li

Generating speech across different accents while preserving speaker identity is crucial for various real-world applications. However, accurately and independently modeling both speaker and accent characteristics in text-to-speech (TTS)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Xuehao Zhou , Mingyang Zhang , Yi Zhou , Zhizheng Wu , Haizhou Li

Recently, the effectiveness of text-to-speech (TTS) systems combined with neural vocoders to generate high-fidelity speech has been shown. However, collecting the required training data and building these advanced systems from scratch are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Yi-Chiao Wu , Patrick Lumban Tobing , Kazuki Yasuhara , Noriyuki Matsunaga , Yamato Ohtani , Tomoki Toda

Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both approaches leave models with little understanding of global…

Accent plays a significant role in speech communication, influencing one's capability to understand as well as conveying a person's identity. This paper introduces a novel and efficient framework for accented Text-to-Speech (TTS) synthesis…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-01 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

In voice conversion (VC), an approach showing promising results in the latest voice conversion challenge (VCC) 2020 is to first use an automatic speech recognition (ASR) model to transcribe the source speech into the underlying linguistic…

Sound · Computer Science 2021-07-21 Wen-Chin Huang , Tomoki Hayashi , Xinjian Li , Shinji Watanabe , Tomoki Toda

Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these…

Sound · Computer Science 2025-10-13 Huu Tuong Tu , Huan Vu , cuong tien nguyen , Dien Hy Ngo , Nguyen Thi Thu Trang

Voice conversion (VC) aims to modify the speaker's timbre while retaining speech content. Previous approaches have tokenized the outputs from self-supervised into semantic tokens, facilitating disentanglement of speech content information.…

Sound · Computer Science 2024-09-11 Zhengyang Chen , Shuai Wang , Mingyang Zhang , Xuechen Liu , Junichi Yamagishi , Yanmin Qian

We propose a text-to-talking-face synthesis framework leveraging latent speech representations from HierSpeech++. A Text-to-Vec module generates Wav2Vec2 embeddings from text, which jointly condition speech and face generation. To handle…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Dogucan Yaman , Seymanur Akti , Fevziye Irem Eyiokur , Alexander Waibel

While neural methods for text-to-speech (TTS) have shown great advances in modeling multiple speakers, even in zero-shot settings, the amount of data needed for those approaches is generally not feasible for the vast majority of the world's…

Computation and Language · Computer Science 2022-10-25 Florian Lux , Julia Koch , Ngoc Thang Vu

The gender of any voice user interface is a key element of its perceived identity. Recently, there has been increasing interest in interfaces where the gender is ambiguous rather than clearly identifying as female or male. This work…