中文
相关论文

相关论文: PRESENT: Zero-Shot Text-to-Prosody Control

200 篇论文

Recent advancements in end-to-end speech synthesis have made it possible to generate highly natural speech. However, training these models typically requires a large amount of high-fidelity speech data, and for unseen texts, the prosody of…

计算与语言 · 计算机科学 2021-11-16 Zhu Li , Yuqing Zhang , Mengxi Nie , Ming Yan , Mengnan He , Ruixiong Zhang , Caixia Gong

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms…

音频与语音处理 · 电气工程与系统科学 2024-04-11 Ziyue Jiang , Jinglin Liu , Yi Ren , Jinzheng He , Zhenhui Ye , Shengpeng Ji , Qian Yang , Chen Zhang , Pengfei Wei , Chunfeng Wang , Xiang Yin , Zejun Ma , Zhou Zhao

End-to-end text-to-speech synthesis systems achieved immense success in recent times, with improved naturalness and intelligibility. However, the end-to-end models, which primarily depend on the attention-based alignment, do not offer an…

音频与语音处理 · 电气工程与系统科学 2021-10-07 Giridhar Pamisetty , K. Sri Rama Murty

This paper proposes a zero-shot text-to-speech (TTS) conditioned by a self-supervised speech-representation model acquired through self-supervised learning (SSL). Conventional methods with embedding vectors from x-vector or global style…

声音 · 计算机科学 2023-12-19 Kenichi Fujita , Takanori Ashihara , Hiroki Kanagawa , Takafumi Moriya , Yusuke Ijima

Generating speech from a face image is crucial for developing virtual humans capable of interacting using their unique voices, without relying on pre-recorded human speech. In this paper, we propose Face-StyleSpeech, a zero-shot…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Minki Kang , Wooseok Han , Eunho Yang

Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional TTS for seen…

声音 · 计算机科学 2023-05-24 Minki Kang , Wooseok Han , Sung Ju Hwang , Eunho Yang

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this…

Modern sequence to sequence neural TTS systems provide close to natural speech quality. Such systems usually comprise a network converting linguistic/phonetic features sequence to an acoustic features sequence, cascaded with a neural…

音频与语音处理 · 电气工程与系统科学 2019-09-26 Slava Shechtman , Alex Sorin

We introduce StyleFusion-TTS, a prompt and/or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research literature. We propose a…

音频与语音处理 · 电气工程与系统科学 2024-09-25 Zhiyong Chen , Xinnuo Li , Zhiqi Ai , Shugong Xu

While controllable Text-to-Speech (TTS) has achieved notable progress, most existing methods remain limited to inter-utterance-level control, making fine-grained intra-utterance expression challenging due to their reliance on non-public…

声音 · 计算机科学 2026-05-19 Qifan Liang , Yuansen Liu , Ruixin Wei , Nan Lu , Junchuan Zhao , Ye Wang

Given a piece of speech and its transcript text, text-based speech editing aims to generate speech that can be seamlessly inserted into the given speech by editing the transcript. Existing methods adopt a two-stage approach: synthesize the…

声音 · 计算机科学 2021-09-14 Chuanxin Tang , Chong Luo , Zhiyuan Zhao , Dacheng Yin , Yucheng Zhao , Wenjun Zeng

Zero-shot emotion transfer in cross-lingual speech synthesis aims to transfer emotion from an arbitrary speech reference in the source language to the synthetic speech in the target language. Building such a system faces challenges of…

声音 · 计算机科学 2023-10-09 Yuke Li , Xinfa Zhu , Yi Lei , Hai Li , Junhui Liu , Danming Xie , Lei Xie

The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often face issues such as slow inference speeds, reliance on complex…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Yinghao Aaron Li , Xilin Jiang , Cong Han , Nima Mesgarani

Recent neural speech synthesis systems have gradually focused on the control of prosody to improve the quality of synthesized speech, but they rarely consider the variability of prosody and the correlation between prosody and semantics…

音频与语音处理 · 电气工程与系统科学 2020-08-14 Zhen Zeng , Jianzong Wang , Ning Cheng , Jing Xiao

While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses…

音频与语音处理 · 电气工程与系统科学 2026-01-13 Tianrui Wang , Haoyu Wang , Meng Ge , Cheng Gong , Chunyu Qiang , Ziyang Ma , Zikang Huang , Guanrou Yang , Xiaobao Wang , Eng Siong Chng , Xie Chen , Longbiao Wang , Jianwu Dang

Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edited speech that is similar to the original speech in terms…

声音 · 计算机科学 2022-10-31 Jason Fong , Yun Wang , Prabhav Agrawal , Vimal Manohar , Jilong Wu , Thilo Köhler , Qing He

Modern zero-shot text-to-speech (TTS) models offer unprecedented expressivity but also pose serious crime risks, as they can synthesize voices of individuals who never consented. In this context, speaker unlearning aims to prevent the…

音频与语音处理 · 电气工程与系统科学 2026-01-29 Myungjin Lee , Eunji Shin , Jiyoung Lee

Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker's voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language…

Speech synthesis has significantly advanced from statistical methods to deep neural network architectures, leading to various text-to-speech (TTS) models that closely mimic human speech patterns. However, capturing nuances such as emotion…

声音 · 计算机科学 2025-01-14 Shaozuo Zhang , Ambuj Mehrish , Yingting Li , Soujanya Poria

Recent research in zero-shot speech synthesis has made significant progress in speaker similarity. However, current efforts focus on timbre generalization rather than prosody modeling, which results in limited naturalness and…

声音 · 计算机科学 2024-06-12 Yuepeng Jiang , Tao Li , Fengyu Yang , Lei Xie , Meng Meng , Yujun Wang
‹ 上一页 1 2 3 10 下一页 ›