English
Related papers

Related papers: VoXtream: Full-Stream Text-to-Speech with Extremel…

200 papers

Text-to-speech (TTS) systems are an important component in voice-based e-commerce applications. These applications include end-to-end voice assistant and customer experience (CX) voice bot. Code-mixed TTS is also relevant in these…

Machine Learning · Computer Science 2023-12-05 Raviraj Joshi , Nikesh Garera

In this paper, we introduce a zero-shot Voice Transfer (VT) module that can be seamlessly integrated into a multi-lingual Text-to-speech (TTS) system to transfer an individual's voice across languages. Our proposed VT module comprises a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Fadi Biadsy , Youzheng Chen , Isaac Elias , Kyle Kastner , Gary Wang , Andrew Rosenberg , Bhuvana Ramabhadran

In this paper we propose Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis with control over speech variation and style transfer. Flowtron borrows insights from IAF and revamps Tacotron in order to…

Sound · Computer Science 2020-07-17 Rafael Valle , Kevin Shih , Ryan Prenger , Bryan Catanzaro

We introduce a text-to-speech (TTS) model called BASE TTS, which stands for $\textbf{B}$ig $\textbf{A}$daptive $\textbf{S}$treamable TTS with $\textbf{E}$mergent abilities. BASE TTS is the largest TTS model to-date, trained on 100K hours of…

Diffusion-based generative models have greatly impacted the speech processing field in recent years, exhibiting high speech naturalness and spawning a new research direction. Their application in real-time communication is, however, still…

Signal Processing · Electrical Eng. & Systems 2026-04-22 Simon Welker , Bunlong Lay , Maris Hillemann , Tal Peer , Timo Gerkmann

Adaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen speakers have diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-04 Yihan Wu , Xu Tan , Bohan Li , Lei He , Sheng Zhao , Ruihua Song , Tao Qin , Tie-Yan Liu

Existing Large Language Model (LLM) based autoregressive (AR) text-to-speech (TTS) systems, while achieving state-of-the-art quality, still face critical challenges. The foundation of this LLM-based paradigm is the discretization of the…

In this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker's voice and enabling arbitrary control and adjustment of speaking style. Prior zero-shot TTS models only mimic the speaker's voice…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-05 Shengpeng Ji , Qian Chen , Wen Wang , Jialong Zuo , Minghui Fang , Ziyue Jiang , Hai Huang , Zehan Wang , Xize Cheng , Siqi Zheng , Zhou Zhao

Spoken dialogue systems often rely on cascaded pipelines that transcribe, process, and resynthesize speech. While effective, this design discards paralinguistic cues and limits expressivity. Recent end-to-end methods reduce latency and…

Whisper is one of the recent state-of-the-art multilingual speech recognition and translation models, however, it is not designed for real time transcription. In this paper, we build on top of Whisper and create Whisper-Streaming, an…

Computation and Language · Computer Science 2023-09-22 Dominik Macháček , Raj Dabre , Ondřej Bojar

End-to-end text-to-speech (TTS) has shown great success on large quantities of paired text plus speech data. However, laborious data collection remains difficult for at least 95% of the languages over the world, which hinders the…

Computation and Language · Computer Science 2019-07-03 Tao Tu , Yuan-Jui Chen , Cheng-chieh Yeh , Hung-yi Lee

The long speech sequence has been troubling language models (LM) based TTS approaches in terms of modeling complexity and efficiency. This work proposes SoCodec, a semantic-ordered multi-stream speech codec, to address this issue. It…

Sound · Computer Science 2024-09-04 Haohan Guo , Fenglong Xie , Kun Xie , Dongchao Yang , Dake Guo , Xixin Wu , Helen Meng

One-shot voice conversion (VC) aims to convert speech from any source speaker to an arbitrary target speaker with only a few seconds of reference speech from the target speaker. This relies heavily on disentangling the speaker's identity…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-02 Yinghao Aaron Li , Cong Han , Nima Mesgarani

We propose LiveGesture, the first fully streamable, speech-driven full-body gesture generation framework that operates with zero look-ahead and supports arbitrary sequence length. Unlike existing co-speech gesture methods, which are…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Muhammad Usama Saleem , Mayur Jagdishbhai Patel , Ekkasit Pinyoanuntapong , Zhongxing Qin , Li Yang , Hongfei Xue , Ahmed Helmy , Chen Chen , Pu Wang

We present a speaker conditioned text-to-speech (TTS) system aimed at addressing challenges in generating speech for unseen speakers and supporting diverse Indian languages. Our method leverages a diffusion-based TTS architecture, where a…

Large language models (LLM)-based speech synthesis has been widely adopted in zero-shot speech synthesis. However, they require a large-scale data and possess the same limitations as previous autoregressive speech models, including slow…

Sound · Computer Science 2023-11-28 Sang-Hoon Lee , Ha-Yeong Choi , Seung-Bin Kim , Seong-Whan Lee

In real-time speech synthesis, neural vocoders often require low-latency synthesis through causal processing and streaming. However, streaming introduces inefficiencies absent in batch synthesis, such as limited parallelism, inter-frame…

Sound · Computer Science 2025-06-05 Reo Yoneyama , Masaya Kawamura , Ryo Terashima , Ryuichi Yamamoto , Tomoki Toda

Traditional text-to-speech (TTS) methods primarily focus on establishing a mapping between phonemes and mel-spectrograms. However, during the phoneme encoding stage, there is often a lack of real mel-spectrogram auxiliary information, which…

Sound · Computer Science 2025-03-11 Tianyun Liu

With the development of automatic speech recognition (ASR) and text-to-speech (TTS) technology, high-quality voice conversion (VC) can be achieved by extracting source content information and target speaker information to reconstruct…

Sound · Computer Science 2023-02-24 Houjian Guo , Chaoran Liu , Carlos Toshinori Ishi , Hiroshi Ishiguro

Automatic Speech Recognition (ASR) plays a crucial role in voice-based applications. For applications requiring real-time feedback like Voice Search, streaming capability becomes vital. While LSTM/RNN and CTC based ASR systems are commonly…

Sound · Computer Science 2023-05-31 Abhinav Goyal , Nikesh Garera
‹ Prev 1 8 9 10 Next ›