English
Related papers

Related papers: Cross-lingual Multispeaker Text-to-Speech under Li…

200 papers

Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both approaches leave models with little understanding of global…

The advancement of multimodal large language models has accelerated the development of speech-to-speech interaction systems. While natural monolingual interaction has been achieved, we find existing models exhibit deficiencies in language…

Computation and Language · Computer Science 2025-10-10 Heyang Liu , Yuhao Wang , Ziyang Cheng , Ronghua Wu , Qunshan Gu , Yanfeng Wang , Yu Wang

Text-to-speech (TTS) technology has achieved impressive results for widely spoken languages, yet many under-resourced languages remain challenged by limited data and linguistic complexities. In this paper, we present a novel methodology…

Sound · Computer Science 2025-04-11 Yizhong Geng , Jizhuo Xu , Zeyu Liang , Jinghan Yang , Xiaoyi Shi , Xiaoyu Shen

This paper proposes a textless training method for many-to-many multilingual speech-to-speech translation that can also benefit the transfer of pre-trained knowledge to text-based systems, text-to-speech synthesis and text-to-speech…

Computation and Language · Computer Science 2024-08-20 Minsu Kim , Jeongsoo Choi , Dahun Kim , Yong Man Ro

The effects of language mismatch impact speech anti-spoofing systems, while investigations and quantification of these effects remain limited. Existing anti-spoofing datasets are mainly in English, and the high cost of acquiring…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-22 Tianchi Liu , Ivan Kukanov , Zihan Pan , Qiongqiong Wang , Hardik B. Sailor , Kong Aik Lee

Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data. In this paper, we closely examine the reasons behind this phenomenon, focusing on the pre-training…

Computation and Language · Computer Science 2025-04-23 Zhijun Wang , Jiahuan Li , Hao Zhou , Rongxiang Weng , Jingang Wang , Xin Huang , Xue Han , Junlan Feng , Chao Deng , Shujian Huang

Language diversity presents a significant challenge in speech-to-text (S2T) tasks, such as automatic speech recognition and translation. Traditional multi-lingual multi-task training approaches aim to address this by jointly optimising…

Sound · Computer Science 2025-07-09 Qiuming Zhao , Guangzhi Sun , Chao Zhang

Deep learning based text-to-speech (TTS) systems have been evolving rapidly with advances in model architectures, training methodologies, and generalization across speakers and languages. However, these advances have not been thoroughly…

Computation and Language · Computer Science 2023-02-20 Gokul Karthik Kumar , Praveen S , Pratyush Kumar , Mitesh M. Khapra , Karthik Nandakumar

This paper aims to build a multi-speaker expressive TTS system, synthesizing a target speaker's speech with multiple styles and emotions. To this end, we propose a novel contrastive learning-based TTS approach to transfer style and emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-26 Xinfa Zhu , Yuke Li , Yi Lei , Ning Jiang , Guoqing Zhao , Lei Xie

We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech--two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks,…

This paper describes progress towards making a Neural Text-to-Speech (TTS) Frontend that works for many languages and can be easily extended to new languages. We take a Machine Translation (MT) inspired approach to constructing the…

Computation and Language · Computer Science 2020-04-13 Alistair Conkie , Andrew Finch

Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the target speaker in the…

Computation and Language · Computer Science 2022-07-14 Yookyung Shin , Younggun Lee , Suhee Jo , Yeongtae Hwang , Taesu Kim

Text-to-Speech (TTS) models can generate natural, human-like speech across multiple languages by transforming phonemes into waveforms. However, multilingual TTS remains challenging due to discrepancies in phoneme vocabularies and variations…

Sound · Computer Science 2025-04-14 Haowei Lou , Hye-young Paik , Sheng Li , Wen Hu , Lina Yao

Cross-lingual voice conversion (VC) is an important and challenging problem due to significant mismatches of the phonetic set and the speech prosody of different languages. In this paper, we build upon the neural text-to-speech (TTS) model,…

Sound · Computer Science 2021-02-04 Shengkui Zhao , Hao Wang , Trung Hieu Nguyen , Bin Ma

Personalizing a speech synthesis system is a highly desired application, where the system can generate speech with the user's voice with rare enrolled recordings. There are two main approaches to build such a system in recent works: speaker…

Sound · Computer Science 2022-08-01 Sung-Feng Huang , Chyi-Jiunn Lin , Da-Rong Liu , Yi-Chen Chen , Hung-yi Lee

Code-switching and language identification in child-directed scenarios present significant challenges, particularly in bilingual environments. This paper addresses this challenge by using Zipformer to handle the nuances of speech, which…

Computation and Language · Computer Science 2025-08-14 Lavanya Shankar , Leibny Paola Garcia Perera

With the rapid development of large language models, researchers have created increasingly advanced spoken dialogue systems that can naturally converse with humans. However, these systems still struggle to handle the full complexity of…

Computation and Language · Computer Science 2025-01-03 Xize Cheng , Dongjie Fu , Xiaoda Yang , Minghui Fang , Ruofan Hu , Jingyu Lu , Bai Jionghao , Zehan Wang , Shengpeng Ji , Rongjie Huang , Linjun Li , Yu Chen , Tao Jin , Zhou Zhao

End-to-end (E2E) systems synthesise high-quality speech, but this typically requires a large amount of data. As E2E synthesis progressed from Tacotron to FastSpeech2, it became evident that features representing prosody, particularly…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-19 Anusha Prakash , S Umesh , Hema A Murthy

End-to-end speech-to-intent classification has shown its advantage in harvesting information from both text and speech. In this paper, we study a technique to develop such an end-to-end system that supports multiple languages. To overcome…

Computation and Language · Computer Science 2021-09-29 Bidisha Sharma , Maulik Madhavi , Xuehao Zhou , Haizhou Li

Multilingual models have been widely used for cross-lingual transfer to low-resource languages. However, the performance on these languages is hindered by their underrepresentation in the pretraining data. To alleviate this problem, we…

Computation and Language · Computer Science 2023-05-29 Tomasz Limisiewicz , Dan Malkin , Gabriel Stanovsky