中文
相关论文

相关论文: Enhancing Emotional Text-to-Speech Controllability…

200 篇论文

Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Yuchi Wang , Junliang Guo , Jianhong Bai , Runyi Yu , Tianyu He , Xu Tan , Xu Sun , Jiang Bian

This paper proposes a neural sequence-to-sequence text-to-speech (TTS) model which can control latent attributes in the generated speech that are rarely annotated in the training data, such as speaking style, accent, background noise, and…

Electromyography-to-Speech (ETS) conversion has demonstrated its potential for silent speech interfaces by generating audible speech from Electromyography (EMG) signals during silent articulations. ETS models usually consist of an EMG…

声音 · 计算机科学 2024-05-15 Zhao Ren , Kevin Scheck , Qinhan Hou , Stefano van Gogh , Michael Wand , Tanja Schultz

Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model based on a conditional…

声音 · 计算机科学 2024-12-05 Jiaxuan Liu , Zhaoci Liu , Yajun Hu , Yingying Gao , Shilei Zhang , Zhenhua Ling

Improving text representation has attracted much attention to achieve expressive text-to-speech (TTS). However, existing works only implicitly learn the prosody with masked token reconstruction tasks, which leads to low training efficiency…

声音 · 计算机科学 2023-05-19 Zhenhui Ye , Rongjie Huang , Yi Ren , Ziyue Jiang , Jinglin Liu , Jinzheng He , Xiang Yin , Zhou Zhao

Expressive Text-to-Speech (TTS) using reference speech has been studied extensively to synthesize natural speech, but there are limitations to obtaining well-represented styles and improving model generalization ability. In this study, we…

音频与语音处理 · 电气工程与系统科学 2024-06-28 Hyun Joon Park , Jin Sob Kim , Wooseok Shin , Sung Won Han

The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion…

音频与语音处理 · 电气工程与系统科学 2024-12-31 Ashishkumar Gudmalwar , Ishan D. Biyani , Nirmesh Shah , Pankaj Wasnik , Rajiv Ratn Shah

Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Soumya Dutta

Recently, sequence-to-sequence models with attention have been successfully applied in Text-to-speech (TTS). These models can generate near-human speech with a large accurately-transcribed speech corpus. However, preparing such a large…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Haitong Zhang , Yue Lin

Emotional talking-head generation has emerged as a pivotal research area at the intersection of computer vision and multimodal artificial intelligence, with its core value lying in enhancing human-computer interaction through immersive and…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Hanlei Shi , Leyuan Qu , Yu Liu , Di Gao , Yuhua Zheng , Taihao Li

Neural network based end-to-end text to speech (TTS) has significantly improved the quality of synthesized speech. Prominent methods (e.g., Tacotron 2) usually first generate mel-spectrogram from text, and then synthesize speech from the…

计算与语言 · 计算机科学 2019-11-21 Yi Ren , Yangjun Ruan , Xu Tan , Tao Qin , Sheng Zhao , Zhou Zhao , Tie-Yan Liu

Reference-based Text-to-Speech (TTS) models can generate multiple, prosodically-different renditions of the same target text. Such models jointly learn a latent acoustic space during training, which can be sampled from during inference.…

计算与语言 · 计算机科学 2023-09-20 Atli Thor Sigurgeirsson , Simon King

The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustrate pho- netic concepts that are otherwise difficult to…

音频与语音处理 · 电气工程与系统科学 2025-11-10 Frederik Rautenberg , Fritz Seebauer , Jana Wiechmann , Michael Kuhlmann , Petra Wagner , Reinhold Haeb-Umbach

Large Language Models (LLMs) have demonstrated superior abilities in tasks such as chatting, reasoning, and question-answering. However, standard LLMs may ignore crucial paralinguistic information, such as sentiment, emotion, and speaking…

Humans can perceive speakers' characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech (TTS) scholars grounded their…

声音 · 计算机科学 2025-04-17 Tian-Hao Zhang , Jiawei Zhang , Jun Wang , Xinyuan Qian , Xu-Cheng Yin

Controlling text-to-speech (TTS) systems to synthesize speech with the prosodic characteristics expected by users has attracted much attention. To achieve controllability, current studies focus on two main directions: (1) using reference…

声音 · 计算机科学 2025-01-09 Weidong Chen , Shan Yang , Guangzhi Li , Xixin Wu

Transfer tasks in text-to-speech (TTS) synthesis - where one or more aspects of the speech of one set of speakers is transferred to another set of speakers that do not feature these aspects originally - remains a challenging task. One of…

Emotion perception and adaptive expression are fundamental capabilities in human-agent interaction. While recent advances in speech emotion captioning (SEC) have improved fine-grained emotional modeling, existing systems remain limited to…

计算与语言 · 计算机科学 2026-04-30 Shuhao Xu , Yifan Hu , Jingjing Wu , Zhihao Du , Zheng Lian , Rui Liu

Spoken meaning often depends not only on what is said, but also on which word is emphasized. The same sentence can convey correction, contrast, or clarification depending on where emphasis falls. Although modern text-to-speech (TTS) systems…

计算与语言 · 计算机科学 2026-04-14 Arnon Turetzky , Avihu Dekel , Hagai Aronowitz , Ron Hoory , Yossi Adi

Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which may be difficult to…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Yifeng Ma , Suzhen Wang , Yu Ding , Bowen Ma , Tangjie Lv , Changjie Fan , Zhipeng Hu , Zhidong Deng , Xin Yu