中文
相关论文

相关论文: AccentBox: Towards High-Fidelity Zero-Shot Accent …

200 篇论文

Advances in speech representation and large language models have enhanced zero-shot text-to-speech (TTS) performance. However, existing zero-shot TTS models face challenges in capturing the complex correlations between acoustic and semantic…

音频与语音处理 · 电气工程与系统科学 2025-08-29 Jingyuan Xing , Zhipeng Li , Jialong Mai , Xiaofen Xing , Xiangmin Xu

This paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distribution matching distillation. The training of E1 TTS is…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Zhijun Liu , Shuai Wang , Pengcheng Zhu , Mengxiao Bi , Haizhou Li

Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Nevertheless, stability and fidelity remain key challenges,…

声音 · 计算机科学 2025-10-24 Hualei Wang , Na Li , Chuke Wang , Shu Wu , Zhifeng Li , Dong Yu

We introduce a zero-shot video captioning method that employs two frozen networks: the GPT-2 language model and the CLIP image-text matching model. The matching score is used to steer the language model toward generating a sentence that has…

计算机视觉与模式识别 · 计算机科学 2022-07-29 Yoad Tewel , Yoav Shalev , Roy Nadler , Idan Schwartz , Lior Wolf

Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in…

音频与语音处理 · 电气工程与系统科学 2024-01-23 Jae-Sung Bae , Joun Yeop Lee , Ji-Hyun Lee , Seongkyu Mun , Taehwa Kang , Hoon-Young Cho , Chanwoo Kim

Generating speech from a face image is crucial for developing virtual humans capable of interacting using their unique voices, without relying on pre-recorded human speech. In this paper, we propose Face-StyleSpeech, a zero-shot…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Minki Kang , Wooseok Han , Eunho Yang

People change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions. However, most text-to-speech (TTS) systems lack the capability to generate speech with rich…

音频与语音处理 · 电气工程与系统科学 2024-09-18 Haibin Wu , Xiaofei Wang , Sefik Emre Eskimez , Manthan Thakker , Daniel Tompkins , Chung-Hsien Tsai , Canrun Li , Zhen Xiao , Sheng Zhao , Jinyu Li , Naoyuki Kanda

Zero-shot speaker cloning aims to synthesize speech for any target speaker unseen during TTS system building, given only a single speech reference of the speaker at hand. Although more practical in real applications, the current zero-shot…

声音 · 计算机科学 2023-10-09 Tao Li , Zhichao Wang , Xinfa Zhu , Jian Cong , Qiao Tian , Yuping Wang , Lei Xie

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intricately encompasses various attributes (e.g., content,…

We present a speaker conditioned text-to-speech (TTS) system aimed at addressing challenges in generating speech for unseen speakers and supporting diverse Indian languages. Our method leverages a diffusion-based TTS architecture, where a…

Zero-shot text-to-speech models can clone a speaker's timbre from a short reference audio, but they also strongly inherit the speaking style present in the reference. As a result, synthesizing speech with a desired style often requires…

音频与语音处理 · 电气工程与系统科学 2026-04-21 Haitao Li , Chunxiang Jin , Chenglin Li , Wenhao Guan , Zhengxing Huang , Xie Chen

We introduce EzAudio, a text-to-audio (T2A) generation framework designed to produce high-quality, natural-sounding sound effects. Core designs include: (1) We propose EzAudio-DiT, an optimized Diffusion Transformer (DiT) designed for audio…

音频与语音处理 · 电气工程与系统科学 2025-06-23 Jiarui Hai , Yong Xu , Hao Zhang , Chenxing Li , Helin Wang , Mounya Elhilali , Dong Yu

Speaker embedding based zero-shot Text-to-Speech (TTS) systems enable high-quality speech synthesis for unseen speakers using minimal data. However, these systems are vulnerable to adversarial attacks, where an attacker introduces…

音频与语音处理 · 电气工程与系统科学 2025-10-07 Ze Li , Yao Shi , Yunfei Xu , Ming Li

Most Zero-shot Multi-speaker TTS (ZS-TTS) systems support only a single language. Although models like YourTTS, VALL-E X, Mega-TTS 2, and Voicebox explored Multilingual ZS-TTS they are limited to just a few high/medium resource languages,…

音频与语音处理 · 电气工程与系统科学 2024-06-10 Edresson Casanova , Kelly Davis , Eren Gölge , Görkem Göknar , Iulian Gulea , Logan Hart , Aya Aljafari , Joshua Meyer , Reuben Morais , Samuel Olayemi , Julian Weber

We propose FEIM-TTS, an innovative zero-shot text-to-speech (TTS) model that synthesizes emotionally expressive speech, aligned with facial images and modulated by emotion intensity. Leveraging deep learning, FEIM-TTS transcends traditional…

声音 · 计算机科学 2024-09-25 Yunji Chu , Yunseob Shim , Unsang Park

In recent years, automatic speech recognition (ASR) models greatly improved transcription performance both in clean, low noise, acoustic conditions and in reverberant environments. However, all these systems rely on the availability of…

音频与语音处理 · 电气工程与系统科学 2024-09-18 Francesco Nespoli , Daniel Barreda , Patrick A. Naylor

Few-shot speaker adaptation is a specific Text-to-Speech (TTS) system that aims to reproduce a novel speaker's voice with a few training data. While numerous attempts have been made to the few-shot speaker adaptation system, there is still…

音频与语音处理 · 电气工程与系统科学 2021-08-17 Ji-Hoon Kim , Sang-Hoon Lee , Ji-Hyun Lee , Hong-Gyu Jung , Seong-Whan Lee

Direct speech-to-speech translation (S2ST) has gradually become popular as it has many advantages compared with cascade S2ST. However, current research mainly focuses on the accuracy of semantic translation and ignores the speech style…

声音 · 计算机科学 2023-07-26 Kun Song , Yi Ren , Yi Lei , Chunfeng Wang , Kun Wei , Lei Xie , Xiang Yin , Zejun Ma

Speech Translation (ST) is the task of translating speech in one language into text in another language. Traditional cascaded approaches for ST, using Automatic Speech Recognition (ASR) and Machine Translation (MT) systems, are prone to…

计算与语言 · 计算机科学 2021-07-14 Tu Anh Dinh

Although diffusion-based, non-autoregressive text-to-speech (TTS) systems have demonstrated impressive zero-shot synthesis capabilities, their efficacy is still hindered by two key challenges: the difficulty of text-speech alignment…

音频与语音处理 · 电气工程与系统科学 2026-02-06 Chunyat Wu , Jiajun Deng , Zhengxi Liu , Zheqi Dai , Haolin He , Qiuqiang Kong