中文
相关论文

相关论文: Intelli-Z: Toward Intelligible Zero-Shot TTS

200 篇论文

Recent advancements in text-to-speech (TTS) technology have increased demand for personalized audio synthesis. Zero-shot voice cloning, a specialized TTS task, aims to synthesize a target speaker's voice using only a single audio sample and…

声音 · 计算机科学 2025-06-03 Ming Meng , Ziyi Yang , Jian Yang , Zhenjie Su , Yonggui Zhu , Zhaoxin Fan

In this paper, we investigate the semi-supervised joint training of text to speech (TTS) and automatic speech recognition (ASR), where a small amount of paired data and a large amount of unpaired text data are available. Conventional…

声音 · 计算机科学 2022-07-12 Naoki Makishima , Satoshi Suzuki , Atsushi Ando , Ryo Masumura

This letter presents an incremental text-to-speech (TTS) method that performs synthesis in small linguistic units while maintaining the naturalness of output speech. Incremental TTS is generally subject to a trade-off between latency and…

声音 · 计算机科学 2021-05-26 Takaaki Saeki , Shinnosuke Takamichi , Hiroshi Saruwatari

The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, tackles this by inserting reasoning steps only during natural…

计算与语言 · 计算机科学 2026-05-21 Xuan Du , Qiangyu Yan , Wenshuo Li , Borui Jiang , Changming Xiao , Han Shu , Xinghao Chen

The lack of clean speech is a practical challenge to the development of speech enhancement systems, which means that there is an inevitable mismatch between their training criterion and evaluation metric. In response to this unfavorable…

声音 · 计算机科学 2023-05-23 Li-Wei Chen , Yao-Fei Cheng , Hung-Shin Lee , Yu Tsao , Hsin-Min Wang

The cloning of a speaker's voice using an untranscribed reference sample is one of the great advances of modern neural text-to-speech (TTS) methods. Approaches for mimicking the prosody of a transcribed reference audio have also been…

声音 · 计算机科学 2022-10-25 Florian Lux , Julia Koch , Ngoc Thang Vu

Neural network based end-to-end Text-to-Speech (TTS) has greatly improved the quality of synthesized speech. While how to use massive spontaneous speech without transcription efficiently still remains an open problem. In this paper, we…

声音 · 计算机科学 2022-02-07 Dabiao Ma , Yitong Zhang , Meng Li , Feng Ye

Identifying keywords in an open-vocabulary context is crucial for personalizing interactions with smart devices. Previous approaches to open vocabulary keyword spotting dependon a shared embedding space created by audio and text encoders.…

人机交互 · 计算机科学 2024-04-19 Kesavaraj V , Anil Kumar Vuppala

Objective evaluation of synthesized speech is critical for advancing speech generation systems, yet existing metrics for intelligibility and prosody remain limited in scope and weakly correlated with human perception. Word Error Rate (WER)…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Ismail Rasim Ulgen , Zongyang Du , Junchen Lu , Philipp Koehn , Berrak Sisman

Zero-shot multi-speaker text-to-speech (ZSM-TTS) models aim to generate a speech sample with the voice characteristic of an unseen speaker. The main challenge of ZSM-TTS is to increase the overall speaker similarity for unseen speakers. One…

音频与语音处理 · 电气工程与系统科学 2022-12-28 Byoung Jin Choi , Myeonghun Jeong , Joun Yeop Lee , Nam Soo Kim

Advances in automated essay scoring (AES) have traditionally relied on labeled essays, requiring tremendous cost and expertise for their acquisition. Recently, large language models (LLMs) have achieved great success in various tasks, but…

计算与语言 · 计算机科学 2024-10-07 Sanwoo Lee , Yida Cai , Desong Meng , Ziyang Wang , Yunfang Wu

On account of growing demands for personalization, the need for a so-called few-shot TTS system that clones speakers with only a few data is emerging. To address this issue, we propose Attentron, a few-shot TTS model that clones voices of…

音频与语音处理 · 电气工程与系统科学 2020-08-13 Seungwoo Choi , Seungju Han , Dongyoung Kim , Sungjoo Ha

Multi-speaker TTS has to learn both linguistic embedding and text embedding to generate speech of desired linguistic content in desired voice. However, it is unclear which characteristic of speech results from speaker and which part from…

音频与语音处理 · 电气工程与系统科学 2020-06-15 Sunghee Jung , Hoirin Kim

End-to-end neural TTS training has shown improved performance in speech style transfer. However, the improvement is still limited by the training data in both target styles and speakers. Inadequate style transfer performance occurs when the…

声音 · 计算机科学 2021-06-21 Xiaochun An , Frank K. Soong , Lei Xie

Recent advances on prompt-tuning cast few-shot classification tasks as a masked language modeling problem. By wrapping input into a template and using a verbalizer which constructs a mapping between label space and label word space,…

计算与语言 · 计算机科学 2022-01-17 Yinyi Wei , Tong Mo , Yongtao Jiang , Weiping Li , Wen Zhao

Zero-shot text-to-speech models can clone a speaker's timbre from a short reference audio, but they also strongly inherit the speaking style present in the reference. As a result, synthesizing speech with a desired style often requires…

音频与语音处理 · 电气工程与系统科学 2026-04-21 Haitao Li , Chunxiang Jin , Chenglin Li , Wenhao Guan , Zhengxing Huang , Xie Chen

This paper aims to build a multi-speaker expressive TTS system, synthesizing a target speaker's speech with multiple styles and emotions. To this end, we propose a novel contrastive learning-based TTS approach to transfer style and emotion…

音频与语音处理 · 电气工程与系统科学 2024-04-26 Xinfa Zhu , Yuke Li , Yi Lei , Ning Jiang , Guoqing Zhao , Lei Xie

Traditional speech separation and speaker diarization approaches rely on prior knowledge of target speakers or a predetermined number of participants in audio signals. To address these limitations, recent advances focus on developing…

Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains…

声音 · 计算机科学 2024-09-17 Xiaoyu Liu , Xu Li , Joan Serrà , Santiago Pascual

Most speaker verification tasks are studied as an open-set evaluation scenario considering the real-world condition. Thus, the generalization power to unseen speakers is of paramount important to the performance of the speaker verification…

音频与语音处理 · 电气工程与系统科学 2021-04-15 Ju-ho Kim , Hye-jin Shim , Jee-weon Jung , Ha-Jin Yu