English
Related papers

Related papers: A Unit-based System and Dataset for Expressive Dir…

200 papers

ESPnet-ST-v2 is a revamp of the open-source ESPnet-ST toolkit necessitated by the broadening interests of the spoken language translation community. ESPnet-ST-v2 supports 1) offline speech-to-text translation (ST), 2) simultaneous…

This paper proposes a direct text to speech translation system using discrete acoustic units. This framework employs text in different source languages as input to generate speech in the target language without the need for text…

Computation and Language · Computer Science 2023-09-15 Victoria Mingote , Pablo Gimeno , Luis Vicente , Sameer Khurana , Antoine Laurent , Jarod Duret

Joint speech-language training is challenging due to the large demand for training data and GPU consumption, as well as the modality gap between speech and language. We present ComSL, a speech-language model built atop a composite…

Computation and Language · Computer Science 2023-10-17 Chenyang Le , Yao Qian , Long Zhou , Shujie Liu , Yanmin Qian , Michael Zeng , Xuedong Huang

End-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance drops substantially.…

Computation and Language · Computer Science 2022-06-10 Biao Zhang , Barry Haddow , Rico Sennrich

Simultaneous speech translation (SST) aims to provide real-time translation of spoken language, even before the speaker finishes their sentence. Traditionally, SST has been addressed primarily by cascaded systems that decompose the task…

Computation and Language · Computer Science 2023-10-18 Peter Polák

End-to-end speech-to-speech translation (S2ST) systems typically struggle with a critical data bottleneck: the scarcity of parallel speech-to-speech corpora. To overcome this, we introduce RosettaSpeech, a novel zero-shot framework trained…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-17 Zhisheng Zheng , Xiaohang Sun , Tuan Dinh , Abhishek Yanamandra , Abhinav Jain , Zhu Liu , Sunil Hadap , Vimal Bhat , Manoj Aggarwal , Gerard Medioni , David Harwath

Recent advancements in end-to-end speech synthesis have made it possible to generate highly natural speech. However, training these models typically requires a large amount of high-fidelity speech data, and for unseen texts, the prosody of…

Computation and Language · Computer Science 2021-11-16 Zhu Li , Yuqing Zhang , Mengxi Nie , Ming Yan , Mengnan He , Ruixiong Zhang , Caixia Gong

Nowadays, training end-to-end neural models for spoken language translation (SLT) still has to confront with extreme data scarcity conditions. The existing SLT parallel corpora are indeed orders of magnitude smaller than those available for…

Computation and Language · Computer Science 2019-10-09 Mattia Antonino Di Gangi , Matteo Negri , Marco Turchi

Empathetic interaction is a cornerstone of human-machine communication, due to the need for understanding speech enriched with paralinguistic cues and generating emotional and expressive responses. However, the most powerful empathetic…

Computation and Language · Computer Science 2025-10-28 Chen Wang , Tianyu Peng , Wen Yang , Yinan Bai , Guangfu Wang , Jun Lin , Lanpeng Jia , Lingxiang Wu , Jinqiao Wang , Chengqing Zong , Jiajun Zhang

End-to-end speech translation models have become a new trend in research due to their potential of reducing error propagation. However, these models still suffer from the challenge of data scarcity. How to effectively use unlabeled or other…

Computation and Language · Computer Science 2021-06-21 Rong Ye , Mingxuan Wang , Lei Li

Speech-to-speech translation is a typical sequence-to-sequence learning task that naturally has two directions. How to effectively leverage bidirectional supervision signals to produce high-fidelity audio for both directions? Existing…

Computation and Language · Computer Science 2023-05-23 Xianchao Wu

Generating expressive and contextually appropriate prosody remains a challenge for modern text-to-speech (TTS) systems. This is particularly evident for long, multi-sentence inputs. In this paper, we examine simple extensions to a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-30 Peter Makarov , Ammar Abbas , Mateusz Łajszczak , Arnaud Joly , Sri Karlapati , Alexis Moinet , Thomas Drugman , Penny Karanasou

We present a neural text-to-speech system for fine-grained prosody transfer from one speaker to another. Conventional approaches for end-to-end prosody transfer typically use either fixed-dimensional or variable-length prosody embedding via…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-05 Viacheslav Klimkov , Srikanth Ronanki , Jonas Rohnke , Thomas Drugman

Compositional speech-to-speech translation (S2ST) systems built upon speech large language models (SpeechLLMs) have recently shown promising performance. However, existing S2ST systems often either neglect source-language information or…

Computation and Language · Computer Science 2026-05-18 Yu Pan , Yang Hou , Xiongfei Wu , Liang Zhang , Yves Le Traon , Lei Ma , Jianjun Zhao

Speech-to-text translation pertains to the task of converting speech signals in a language to text in another language. It finds its application in various domains, such as hands-free communication, dictation, video lecture transcription,…

Computation and Language · Computer Science 2024-06-11 Nivedita Sethiya , Chandresh Kumar Maurya

We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting conversational…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-11 Manuel Sam Ribeiro , Julian Roth , Giulia Comini , Goeric Huybrechts , Adam Gabrys , Jaime Lorenzo-Trueba

This research investigates the Statistical Machine Translation approaches to translate speech in real time automatically. Such systems can be used in a pipeline with speech recognition and synthesis software in order to produce a real-time…

Computation and Language · Computer Science 2015-10-01 Krzysztof Wołk , Krzysztof Marasek

This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation.…

Computation and Language · Computer Science 2024-01-17 Eliya Nachmani , Alon Levkovitch , Yifan Ding , Chulayuth Asawaroengchai , Heiga Zen , Michelle Tadmor Ramanovich

We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework. ESPnet-ST is a new project inside end-to-end speech processing toolkit, ESPnet, which integrates or newly…

Computation and Language · Computer Science 2020-10-01 Hirofumi Inaguma , Shun Kiyono , Kevin Duh , Shigeki Karita , Nelson Enrique Yalta Soplin , Tomoki Hayashi , Shinji Watanabe

Speech-to-Speech (S2S) models have shown promising dialogue capabilities, but their ability to handle paralinguistic cues - such as emotion, tone, and speaker attributes - and to respond appropriately in both content and style remains…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-09 Shu-wen Yang , Ming Tu , Andy T. Liu , Xinghua Qu , Hung-yi Lee , Lu Lu , Yuxuan Wang , Yonghui Wu
‹ Prev 1 3 4 5 6 7 10 Next ›