English
Related papers

Related papers: FCTalker: Fine and Coarse Grained Context Modeling…

200 papers

Speech synthesis technology has witnessed significant advancements in recent years, enabling the creation of natural and expressive synthetic speech. One area of particular interest is the generation of synthetic child speech, which…

Sound · Computer Science 2023-11-09 Rishabh Jain , Peter Corcoran

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-art end-to-end text-to-speech synthesis system, to uncover…

Computation and Language · Computer Science 2018-08-07 Daisy Stanton , Yuxuan Wang , RJ Skerry-Ryan

Recently, denoising diffusion probabilistic models and generative score matching have shown high potential in modelling complex data distributions while stochastic calculus has provided a unified point of view on these techniques allowing…

Machine Learning · Computer Science 2021-08-06 Vadim Popov , Ivan Vovk , Vladimir Gogoryan , Tasnima Sadekova , Mikhail Kudinov

Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional…

Sound · Computer Science 2024-12-13 Weizhen Bian , Yubo Zhou , Kaitai Zhang , Xiaohan Gu

Despite imperfect score-matching causing drift in training and sampling distributions of diffusion models, recent advances in diffusion-based acoustic models have revolutionized data-sufficient single-speaker Text-to-Speech (TTS)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-01 Heyang Xue , Shuai Guo , Pengcheng Zhu , Mengxiao Bi

Document-level discourse parsing, in accordance with the Rhetorical Structure Theory (RST), remains notoriously challenging. Challenges include the deep structure of document-level discourse trees, the requirement of subtle semantic…

Computation and Language · Computer Science 2020-12-22 Ke Shi , Zhengyuan Liu , Nancy F. Chen

This chapter presents a novel approach to brain-to-speech (BTS) synthesis from intracranial electroencephalography (iEEG) data, emphasizing prosody-aware feature engineering and advanced transformer-based models for high-fidelity speech…

Signal Processing · Electrical Eng. & Systems 2026-04-08 Mohammed Salah Al-Radhi , Géza Németh , Andon Tchechmedjiev , Binbin Xu

This paper presents a method to accelerate the inference process of diffusion transformer (DiT)-based text-to-speech (TTS) models by applying a selective caching mechanism to transformer layers. Specifically, I integrate SmoothCache into…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-11 Siratish Sakpiboonchit

For text-to-speech (TTS) synthesis, prosodic structure prediction (PSP) plays an important role in producing natural and intelligible speech. Although inter-utterance linguistic information can influence the speech interpretation of the…

Sound · Computer Science 2023-09-01 Jie Chen , Changhe Song , Deyi Tuo , Xixin Wu , Shiyin Kang , Zhiyong Wu , Helen Meng

Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody. We introduce ED-TTS, a multi-scale emotional speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-17 Haobin Tang , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

Generating speech across different accents while preserving speaker identity is crucial for various real-world applications. However, accurately and independently modeling both speaker and accent characteristics in text-to-speech (TTS)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Xuehao Zhou , Mingyang Zhang , Yi Zhou , Zhizheng Wu , Haizhou Li

This paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrary-length speech insertion and even full sentence generation. In the proposed paradigm, global and local factors in…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-29 Dacheng Yin , Chuanxin Tang , Yanqing Liu , Xiaoqiang Wang , Zhiyuan Zhao , Yucheng Zhao , Zhiwei Xiong , Sheng Zhao , Chong Luo

To further improve the speaking styles of synthesized speeches, current text-to-speech (TTS) synthesis systems commonly employ reference speeches to stylize their outputs instead of just the input texts. These reference speeches are…

Sound · Computer Science 2023-08-31 Yi Meng , Xiang Li , Zhiyong Wu , Tingtian Li , Zixun Sun , Xinyu Xiao , Chi Sun , Hui Zhan , Helen Meng

As recent text-to-speech (TTS) systems have been rapidly improved in speech quality and generation speed, many researchers now focus on a more challenging issue: expressive TTS. To control speaking styles, existing expressive TTS models use…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-07 Minchan Kim , Sung Jun Cheon , Byoung Jin Choi , Jong Jin Kim , Nam Soo Kim

Novice content creators often invest significant time recording expressive speech for social media videos. While recent advancements in text-to-speech (TTS) technology can generate highly realistic speech in various languages and accents,…

Human-Computer Interaction · Computer Science 2025-04-08 Stephen Brade , Sam Anderson , Rithesh Kumar , Zeyu Jin , Anh Truong

Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of labeled data, which is…

Sound · Computer Science 2022-06-28 Yihan Wu , Xi Wang , Shaofei Zhang , Lei He , Ruihua Song , Jian-Yun Nie

Prosody is essential for speech technology, shaping comprehension, naturalness, and expressiveness. However, current text-to-speech (TTS) systems still struggle to accurately capture human-like prosodic variation, in part because existing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-05 Cedric Chan , Jianjing Kuang

This study proposes a segmental-level prosodic probing framework to evaluate neural TTS models' ability to reproduce consonant-induced f0 perturbation, a fine-grained segmental-prosodic effect that reflects local articulatory mechanisms. We…

Computation and Language · Computer Science 2026-03-24 Tianle Yang , Chengzhe Sun , Phil Rose , Cassandra L. Jacobs , Siwei Lyu

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to-speech (TTS) screen…

Social and Information Networks · Computer Science 2024-10-28 Suparna De , Ionut Bostan , Nishanth Sastry

We present a neural text-to-speech system for fine-grained prosody transfer from one speaker to another. Conventional approaches for end-to-end prosody transfer typically use either fixed-dimensional or variable-length prosody embedding via…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-05 Viacheslav Klimkov , Srikanth Ronanki , Jonas Rohnke , Thomas Drugman