English
Related papers

Related papers: GTR-Voice: Articulatory Phonetics Informed Control…

200 papers

In multi-speaker speech synthesis, data from a number of speakers usually tend to have great diversity due to the fact that the speakers may differ largely in ages, speaking styles, emotions, and so on. It is important but challenging to…

Sound · Computer Science 2022-02-14 Qinghua Wu , Quanbo Shen , Jian Luan , YuJun Wang

In recent years, prompting has quickly become one of the standard ways of steering the outputs of generative machine learning models, due to its intuitive use of natural language. In this work, we propose a system conditioned on embeddings…

Computation and Language · Computer Science 2024-06-13 Thomas Bott , Florian Lux , Ngoc Thang Vu

Recent work has shown that it is possible to resynthesize high-quality speech based, not on text, but on low bitrate discrete units that have been learned in a self-supervised fashion and can therefore capture expressive aspects of speech…

Large-scale text-to-speech (TTS) models have made significant progress recently.However, they still fall short in the generation of Chinese dialectal speech. Toaddress this, we propose Bailing-TTS, a family of large-scale TTS models capable…

Computation and Language · Computer Science 2024-08-02 Xinhan Di , Zihao Chen , Yunming Liang , Junjie Zheng , Yihua Wang , Chaofan Ding

Expressive speech synthesis is crucial for many human-computer interaction scenarios, such as audiobooks, podcasts, and voice assistants. Previous works focus on predicting the style embeddings at one single scale from the information…

Sound · Computer Science 2023-08-01 Shun Lei , Yixuan Zhou , Liyang Chen , Zhiyong Wu , Xixin Wu , Shiyin Kang , Helen Meng

In this paper, we propose a neural articulation-to-speech (ATS) framework that synthesizes high-quality speech from articulatory signal in a multi-speaker situation. Most conventional ATS approaches only focus on modeling contextual…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Miseul Kim , Zhenyu Piao , Jihyun Lee , Hong-Goo Kang

Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise control over attributes…

Sound · Computer Science 2025-09-10 Yerin Ryu , Inseop Shin , Chanwoo Kim

Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding results a plausible…

Computation and Language · Computer Science 2026-03-18 Yuan Ge , Haishu Zhao , Aokai Hao , Junxiang Zhang , Bei Li , Xiaoqian Liu , Chenglong Wang , Jianjin Wang , Bingsen Zhou , Bingyu Liu , Jingbo Zhu , Zhengtao Yu , Tong Xiao

Singing voice synthesis has made remarkable progress in generating natural and high-quality voices. However, existing methods rarely provide precise control over vocal techniques such as intensity, mixed voice, falsetto, bubble, and breathy…

Sound · Computer Science 2025-04-22 Wenxiang Guo , Yu Zhang , Changhao Pan , Rongjie Huang , Li Tang , Ruiqi Li , Zhiqing Hong , Yongqi Wang , Zhou Zhao

In recent years, the field of image generation has been revolutionized by the application of autoregressive transformers and DDPMs. These approaches model the process of image generation as a step-wise probabilistic processes and leverage…

Sound · Computer Science 2023-05-25 James Betker

Emotional text-to-speech synthesis (ETTS) has seen much progress in recent years. However, the generated voice is often not perceptually identifiable by its intended emotion category. To address this problem, we propose a new interactive…

Computation and Language · Computer Science 2021-06-15 Rui Liu , Berrak Sisman , Haizhou Li

Various parametric representations have been proposed to model the speech signal. While the performance of such vocoders is well-known in the context of speech processing, their extrapolation to singing voice synthesis might not be…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-09 Onur Babacan , Thomas Drugman , Tuomo Raitio , Daniel Erro , Thierry Dutoit

In recent years, there has been significant progress in Text-to-Speech (TTS) synthesis technology, enabling the high-quality synthesis of voices in common scenarios. In unseen situations, adaptive TTS requires a strong generalization…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Zhipeng Li , Xiaofen Xing , Jun Wang , Shuaiqi Chen , Guoqiao Yu , Guanglu Wan , Xiangmin Xu

The purpose of this study is to investigate how humans interpret musical scores expressively, and then design machines that sing like humans. We consider six factors that have a strong influence on the expression of human singing. The…

Sound · Computer Science 2015-02-17 Ju-Chiang Wang , Hung-Yan Gu , Hsin-Min Wang

Generating expressive and controllable human speech is one of the core goals of generative artificial intelligence, but its progress has long been constrained by two fundamental challenges: the deep entanglement of speech factors and the…

Sound · Computer Science 2025-11-20 Xinyue Yu , Youqing Fang , Pingyu Wu , Guoyang Ye , Wenbo Zhou , Weiming Zhang , Song Xiao

This paper introduces EmoSSLSphere, a novel framework for multilingual emotional text-to-speech (TTS) synthesis that combines spherical emotion vectors with discrete token features derived from self-supervised learning (SSL). By encoding…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-07 Joonyong Park , Kenichi Nakamura

Use a parametric representation of audio to train a generative model in the interest of obtaining more flexible control over the generated sound.

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-20 Krishna Subramani , Alexandre D'Hooge , Preeti Rao

When people try to influence others to do something, they subconsciously adjust their speech to include appropriate emotional information. In order for a robot to influence people in the same way, the robot should be able to imitate the…

It is increasingly considered that human speech perception and production both rely on articulatory representations. In this paper, we investigate whether this type of representation could improve the performances of a deep generative model…

Sound · Computer Science 2021-04-08 Marc-Antoine Georges , Laurent Girin , Jean-Luc Schwartz , Thomas Hueber

With the advancement of speech synthesis technology, users have higher expectations for the naturalness and expressiveness of synthesized speech. But previous research ignores the importance of prompt selection. This study proposes a…

Sound · Computer Science 2025-04-15 Dan Luo , Chengyuan Ma , Weiqin Li , Jun Wang , Wei Chen , Zhiyong Wu
‹ Prev 1 4 5 6 7 8 10 Next ›