中文
相关论文

相关论文: Emotion-Aware Speech Generation with Character-Spe…

200 篇论文

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have…

声音 · 计算机科学 2024-10-17 Mithun Manivannan , Vignesh Nethrapalli , Mark Cartwright

Generative models have advanced rapidly, enabling impressive talking head generation that brings AI to life. However, most existing methods focus solely on one-way portrait animation. Even the few that support bidirectional conversational…

音频与语音处理 · 电气工程与系统科学 2025-11-25 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

The increasing use of dialogue agents makes it extremely desirable for them to understand and acknowledge the implied emotions to respond like humans with empathy. Chatbots using traditional techniques analyze emotions based on the context…

计算与语言 · 计算机科学 2021-05-27 Akhilesh Ravi , Amit Yadav , Jainish Chauhan , Jatin Dholakia , Naman Jain , Mayank Singh

Contemporary conversational systems often present a significant limitation: their responses lack the emotional depth and disfluent characteristic of human interactions. This absence becomes particularly noticeable when users seek more…

计算与语言 · 计算机科学 2024-04-03 Rohan Chaudhury , Mihir Godbole , Aakash Garg , Jinsil Hwaryoung Seo

Modern text-to-speech synthesis pipelines typically involve multiple processing stages, each of which is designed or learnt independently from the rest. In this work, we take on the challenging task of learning to synthesise speech from…

声音 · 计算机科学 2021-03-18 Jeff Donahue , Sander Dieleman , Mikołaj Bińkowski , Erich Elsen , Karen Simonyan

In this work, we tackle a problem of speech emotion classification. One of the issues in the area of affective computation is that the amount of annotated data is very limited. On the other hand, the number of ways that the same emotion can…

计算与语言 · 计算机科学 2018-04-02 Egor Lakomkin , Cornelius Weber , Stefan Wermter

This paper investigates a novel task of talking face video generation solely from speeches. The speech-to-video generation technique can spark interesting applications in entertainment, customer service, and human-computer-interaction…

声音 · 计算机科学 2021-07-15 Shijing Si , Jianzong Wang , Xiaoyang Qu , Ning Cheng , Wenqi Wei , Xinghua Zhu , Jing Xiao

Recent studies have outlined the accessibility challenges faced by blind or visually impaired, and less-literate people, in interacting with social networks, in-spite of facilitating technologies such as monotone text-to-speech (TTS) screen…

社会与信息网络 · 计算机科学 2024-10-28 Suparna De , Ionut Bostan , Nishanth Sastry

Story ending generation aims at generating reasonable endings for a given story context. Most existing studies in this area focus on generating coherent or diversified story endings, while they ignore that different characters may lead to…

计算与语言 · 计算机科学 2022-09-02 Xinyu Jiang , Qi Zhang , Chongyang Shi , Kaiying Jiang , Liang Hu , Shoujin Wang

Conversational Speech Synthesis (CSS) aims to align synthesized speech with the emotional and stylistic context of user-agent interactions to achieve empathy. Current generative CSS models face interpretability limitations due to…

声音 · 计算机科学 2025-05-20 Yifan Hu , Rui Liu , Yi Ren , Xiang Yin , Haizhou Li

In an era of human-computer interaction with increasingly agentic AI systems capable of connecting with users conversationally, speech is an important modality for commanding agents. By recognizing and using speech emotions (i.e., how a…

人机交互 · 计算机科学 2025-04-14 Ilhan Aslan , Timothy Merritt , Stine S. Johansen , Niels van Berkel

Empathy, which is widely used in psychological counselling, is a key trait of everyday human conversations. Equipped with commonsense knowledge, current approaches to empathetic response generation focus on capturing implicit emotion within…

计算与语言 · 计算机科学 2022-11-15 Lanrui Wang , Jiangnan Li , Zheng Lin , Fandong Meng , Chenxu Yang , Weiping Wang , Jie Zhou

Centrality of emotion for the stories told by humans is underpinned by numerous studies in literature and psychology. The research in automatic storytelling has recently turned towards emotional storytelling, in which characters' emotions…

计算与语言 · 计算机科学 2019-06-07 Evgeny Kim , Roman Klinger

We present a methodology to train our multi-speaker emotional text-to-speech synthesizer that can express speech for 10 speakers' 7 different emotions. All silences from audio samples are removed prior to learning. This results in fast…

计算与语言 · 计算机科学 2021-12-08 Sungjae Cho , Soo-Young Lee

Explicitly modeling emotions in dialogue generation has important applications, such as building empathetic personal companions. In this study, we consider the task of expressing a specific emotion for dialogue generation. Previous…

计算与语言 · 计算机科学 2021-09-23 Chengzhang Dong , Chenyang Huang , Osmar Zaïane , Lili Mou

Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs focus on generating…

人机交互 · 计算机科学 2025-09-19 Taesoo Kim , Yongsik Jo , Hyunmin Song , Taehwan Kim

The consistency of a response to a given post at semantic-level and emotional-level is essential for a dialogue system to deliver human-like interactions. However, this challenge is not well addressed in the literature, since most of the…

计算与语言 · 计算机科学 2021-04-12 Wei Wei , Jiayi Liu , Xianling Mao , Guibin Guo , Feida Zhu , Pan Zhou , Yuchong Hu , Shanshan Feng

Existing text-to-speech systems predominantly focus on single-sentence synthesis and lack adequate contextual modeling as well as fine-grained performance control capabilities for generating coherent multicast audiobooks. To address these…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Min Liu , JingJing Yin , Xiang Zhang , Siyu Hao , Yanni Hu , Bin Lin , Yuan Feng , Hongbin Zhou , Jianhao Ye

Generating emotional language is a key step towards building empathetic natural language processing agents. However, a major challenge for this line of research is the lack of large-scale labeled training data, and previous studies are…

计算与语言 · 计算机科学 2018-05-15 Xianda Zhou , William Yang Wang

Due to the increasing demand in films and games, synthesizing 3D avatar animation has attracted much attention recently. In this work, we present a production-ready text/speech-driven full-body animation synthesis system. Given the text and…

图形学 · 计算机科学 2022-06-01 Wenlin Zhuang , Jinwei Qi , Peng Zhang , Bang Zhang , Ping Tan