English
Related papers

Related papers: Position: Towards Responsible Evaluation for Text-…

200 papers

In recent years, Text-to-Image (T2I) models have been extensively studied, especially with the emergence of diffusion models that achieve state-of-the-art results on T2I synthesis tasks. However, existing benchmarks heavily rely on…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Eslam Mohamed Bakr , Pengzhan Sun , Xiaoqian Shen , Faizan Farooq Khan , Li Erran Li , Mohamed Elhoseiny

Speech synthesis (text to speech, TTS) and recognition (automatic speech recognition, ASR) are important speech tasks, and require a large amount of text and speech pairs for model training. However, there are more than 6,000 languages in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Jin Xu , Xu Tan , Yi Ren , Tao Qin , Jian Li , Sheng Zhao , Tie-Yan Liu

While recent text-to-speech (TTS) systems increasingly integrate nonverbal vocalizations (NVs), their evaluations lack standardized metrics and reliable ground-truth references. To bridge this gap, we propose NV-Bench, the first benchmark…

Sound · Computer Science 2026-03-19 Qinke Ni , Huan Liao , Dekun Chen , Yuxiang Wang , Zhizheng Wu

Instruction-guided text-to-speech (ITTS) enables users to control speech generation through natural language prompts, offering a more intuitive interface than traditional TTS. However, the alignment between user style instructions and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-28 Yi-Cheng Lin , Huang-Cheng Chou , Tzu-Chieh Wei , Kuan-Yu Chen , Hung-yi Lee

With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve the MIDI-to-audio performance by borrowing from TTS…

Sound · Computer Science 2023-03-22 Xuan Shi , Erica Cooper , Xin Wang , Junichi Yamagishi , Shrikanth Narayanan

State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data, standard TTS models suffer from speech quality and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-17 Adam Gabryś , Goeric Huybrechts , Manuel Sam Ribeiro , Chung-Ming Chien , Julian Roth , Giulia Comini , Roberto Barra-Chicote , Bartek Perz , Jaime Lorenzo-Trueba

While recent text-to-speech (TTS) systems have made remarkable strides toward human-level quality, the performance of cross-lingual TTS lags behind that of intra-lingual TTS. This gap is mainly rooted from the speaker-language entanglement…

Sound · Computer Science 2023-06-13 Ji-Hoon Kim , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim

Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We propose Spotlight-TTS,…

Sound · Computer Science 2025-07-01 Nam-Gyu Kim , Deok-Hyeon Cho , Seung-Bin Kim , Seong-Whan Lee

The evaluation of synthetic and processed speech has long been a cornerstone of audio engineering and speech science. Although subjective listening tests remain the gold standard for assessing perceptual quality and intelligibility, their…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-05 Yu Tsao

In this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human-level TTS synthesis. StyleTTS 2 differs from its…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Yinghao Aaron Li , Cong Han , Vinay S. Raghavan , Gavin Mischler , Nima Mesgarani

Text to speech (TTS) and automatic speech recognition (ASR) are two dual tasks in speech processing and both achieve impressive performance thanks to the recent advance in deep learning and large amount of aligned speech and text data.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-28 Yi Ren , Xu Tan , Tao Qin , Sheng Zhao , Zhou Zhao , Tie-Yan Liu

This report explores the challenge of enhancing expressiveness control in Text-to-Speech (TTS) models by augmenting a frozen pretrained model with a Diffusion Model that is conditioned on joint semantic audio/text embeddings. The paper…

Computation and Language · Computer Science 2023-11-21 Mathias Vogel

Voice directors often iteratively refine voice actors' performances by providing feedback to achieve the desired outcome. While this iterative feedback-based refinement process is important in actual recordings, it has been overlooked in…

Sound · Computer Science 2025-07-03 Hiroki Kanagawa , Kenichi Fujita , Aya Watanabe , Yusuke Ijima

This work proposes FireRedTTS, a foundation text-to-speech framework, to meet the growing demands for personalized and diverse generative speech applications. The framework comprises three parts: data processing, foundation system, and…

Sound · Computer Science 2025-04-14 Hao-Han Guo , Yao Hu , Kun Liu , Fei-Yu Shen , Xu Tang , Yi-Chen Wu , Feng-Long Xie , Kun Xie , Kai-Tuo Xu

Controllable TTS models with natural language prompts often lack the ability for fine-grained control and face a scarcity of high-quality data. We propose a two-stage style-controllable TTS system with language models, utilizing a quantized…

Multimedia · Computer Science 2025-06-04 Yongqi Wang , Chunlei Zhang , Hangting Chen , Zhou Zhao , Dong Yu

Previous multilingual text-to-speech (TTS) approaches have considered leveraging monolingual speaker data to enable cross-lingual speech synthesis. However, such data-efficient approaches have ignored synthesizing emotional aspects of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-01 Xinfa Zhu , Yi Lei , Tao Li , Yongmao Zhang , Hongbin Zhou , Heng Lu , Lei Xie

Modern text simplification (TS) heavily relies on the availability of gold standard data to build machine learning models. However, existing studies show that parallel TS corpora contain inaccurate simplifications and incorrect alignments.…

Computation and Language · Computer Science 2021-07-30 Laura Vásquez-Rodríguez , Matthew Shardlow , Piotr Przybyła , Sophia Ananiadou

Empathy is a vital factor that contributes to mutual understanding, and joint problem-solving. In recent years, a growing number of studies have recognized the benefits of empathy and started to incorporate empathy in conversational…

Computation and Language · Computer Science 2023-10-13 Aravind Sesagiri Raamkumar , Yinping Yang

Traditional Text-to-Speech (TTS) systems rely on studio-quality speech recorded in controlled settings.a Recently, an effort known as noisy-TTS training has emerged, aiming to utilize in-the-wild data. However, the lack of dedicated…

Think about how human handles complex reading tasks: marking key points, inferring their relationships, and structuring information to guide understanding and responses. Likewise, can a large language model benefit from text structure to…

Computation and Language · Computer Science 2026-03-05 Qinsi Wang , Hancheng Ye , Jinhee Kim , Jinghan Ke , Yifei Wang , Martin Kuo , Zishan Shao , Dongting Li , Yueqian Lin , Ting Jiang , Chiyue Wei , Qi Qian , Wei Wen , Helen Li , Yiran Chen
‹ Prev 1 8 9 10 Next ›