中文
相关论文

相关论文: Maximizing Mutual Information for Tacotron

200 篇论文

Synthetic data becomes crucial for large language model training, but its effectiveness is highly inconsistent. We provide an information-theoretic account of this inconsistency: synthetic data improves a model only when the…

机器学习 · 计算机科学 2026-05-19 Hanyu Li , Zhengqi Sun , Xiaotie Deng

Estimation of information theoretic quantities such as mutual information and its conditional variant has drawn interest in recent times owing to their multifaceted applications. Newly proposed neural estimators for these quantities have…

Creating synthetic voices with found data is challenging, as real-world recordings often contain various types of audio degradation. One way to address this problem is to pre-enhance the speech with an enhancement model and then use the…

音频与语音处理 · 电气工程与系统科学 2023-10-03 Yusheng Tian , Wei Liu , Tan Lee

Although many pretrained models exist for text or images, there have been relatively fewer attempts to train representations specifically for dialog understanding. Prior works usually relied on finetuned representations based on generic…

计算与语言 · 计算机科学 2022-05-04 Bishal Santra , Sumegh Roychowdhury , Aishik Mandal , Vasu Gurram , Atharva Naik , Manish Gupta , Pawan Goyal

Accurately finding the wrong words in the automatic speech recognition (ASR) hypothesis and recovering them well-founded is the goal of speech error correction. In this paper, we propose a non-autoregressive speech error correction method.…

计算与语言 · 计算机科学 2024-07-19 Yuchun Shu , Bo Hu , Yifeng He , Hao Shi , Longbiao Wang , Jianwu Dang

Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data…

Regressive Text-to-Speech (TTS) system utilizes attention mechanism to generate alignment between text and acoustic feature sequence. Alignment determines synthesis robustness (e.g, the occurence of skipping, repeating, and collapse) and…

人工智能 · 计算机科学 2023-06-06 Dengfeng Ke , Yayue Deng , Yukang Jia , Jinlong Xue , Qi Luo , Ya Li , Jianqing Sun , Jiaen Liang , Binghuai Lin

Recent advances in maximizing mutual information (MI) between the source and target have demonstrated its effectiveness in text generation. However, previous works paid little attention to modeling the backward network of MI (i.e.,…

计算与语言 · 计算机科学 2020-07-02 Boyuan Pan , Yazheng Yang , Kaizhao Liang , Bhavya Kailkhura , Zhongming Jin , Xian-Sheng Hua , Deng Cai , Bo Li

Tacotron-based end-to-end speech synthesis has shown remarkable voice quality. However, the rendering of prosody in the synthesized speech remains to be improved, especially for long sentences, where prosodic phrasing errors can occur…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Rui Liu , Berrak Sisman , Feilong Bao , Guanglai Gao , Haizhou Li

Recent neural networks such as WaveNet and sampleRNN that learn directly from speech waveform samples have achieved very high-quality synthetic speech in terms of both naturalness and speaker similarity even in multi-speaker text-to-speech…

音频与语音处理 · 电气工程与系统科学 2018-08-01 Yi Zhao , Shinji Takaki , Hieu-Thi Luong , Junichi Yamagishi , Daisuke Saito , Nobuaki Minematsu

We describe a sequence-to-sequence neural network which directly generates speech waveforms from text inputs. The architecture extends the Tacotron model by incorporating a normalizing flow into the autoregressive decoder loop. Output…

计算与语言 · 计算机科学 2021-02-09 Ron J. Weiss , RJ Skerry-Ryan , Eric Battenberg , Soroosh Mariooryad , Diederik P. Kingma

The task of synthetic speech generation is to generate language content from a given text, then simulating fake human voice.The key factors that determine the effect of synthetic speech generation mainly include speed of generation,…

声音 · 计算机科学 2023-07-04 Sheng Zhao , Qilong Yuan , Yibo Duan , Zhuoyue Chen

Autonomous soundscape augmentation systems typically use trained models to pick optimal maskers to effect a desired perceptual change. While acoustic information is paramount to such systems, contextual information, including participant…

声音 · 计算机科学 2024-07-03 Kenneth Ooi , Karn N. Watcharasupat , Bhan Lam , Zhen-Ting Ong , Woon-Seng Gan

Learning disentangled representations of textual data is essential for many natural language tasks such as fair classification, style transfer and sentence generation, among others. The existent dominant approaches in the context of text…

人工智能 · 计算机科学 2021-05-07 Pierre Colombo , Chloe Clavel , Pablo Piantanida

An information-theoretic upper bound on the generalization error of supervised learning algorithms is derived. The bound is constructed in terms of the mutual information between each individual training sample and the output of the…

机器学习 · 计算机科学 2020-08-06 Yuheng Bu , Shaofeng Zou , Venugopal V. Veeravalli

Contemporary neural speech synthesis models have indeed demonstrated remarkable proficiency in synthetic speech generation as they have attained a level of quality comparable to that of human-produced speech. Nevertheless, it is important…

计算与语言 · 计算机科学 2024-04-04 Yejin Jeon , Yunsu Kim , Gary Geunbae Lee

The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding…

声音 · 计算机科学 2021-10-12 Cheng Gong , Longbiao Wang , Zhenhua Ling , Ju Zhang , Jianwu Dang

Aligning text-to-speech (TTS) system outputs with human feedback through preference optimization has been shown to effectively improve the robustness and naturalness of language model-based TTS models. Current approaches primarily require…

计算与语言 · 计算机科学 2026-04-28 Rikuto Kotoge , Yuichi Sasaki

Research on automatic speech recognition (ASR) systems for electrolaryngeal speakers has been relatively unexplored due to small datasets. When training data is lacking in ASR, a large-scale pretraining and fine tuning framework is often…

声音 · 计算机科学 2023-05-31 Lester Phillip Violeta , Ding Ma , Wen-Chin Huang , Tomoki Toda

The GPT-4 technical report suggests that downstream performance can be predicted from pre-training signals, but offers little methodological detail on how to quantify this. This work address this gap by modeling knowledge retention, the…