中文
相关论文

相关论文: Affectron: Emotional Speech Synthesis with Affecti…

200 篇论文

Understanding emotions in natural language is inherently a multi-dimensional reasoning problem, where multiple affective signals interact through context, interpersonal relations, and situational cues. However, most existing emotion…

计算与语言 · 计算机科学 2026-04-02 Hemanth Kotaprolu , Kishan Maharaj , Raey Zhao , Abhijit Mishra , Pushpak Bhattacharyya

While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's speaking status…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Yasheng Sun , Wenqing Chu , Hang Zhou , Kaisiyuan Wang , Hideki Koike

At present emotion extraction from speech is a very important issue due to its diverse applications. Hence, it becomes absolutely necessary to obtain models that take into consideration the speaking styles of a person, vocal tract…

People change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions. However, most text-to-speech (TTS) systems lack the capability to generate speech with rich…

音频与语音处理 · 电气工程与系统科学 2024-09-18 Haibin Wu , Xiaofei Wang , Sefik Emre Eskimez , Manthan Thakker , Daniel Tompkins , Chung-Hsien Tsai , Canrun Li , Zhen Xiao , Sheng Zhao , Jinyu Li , Naoyuki Kanda

In the field of affective computing, traditional methods for generating emotions predominantly rely on deep learning techniques and large-scale emotion datasets. However, deep learning techniques are often complex and difficult to…

人机交互 · 计算机科学 2025-03-24 Haidong Wang , Qia Shan , JianHua Zhang , PengFei Xiao , Ao Liu

The success of automatic speaker verification shows that discriminative speaker representations can be extracted from neutral speech. However, as a kind of non-verbal voice, laughter should also carry speaker information intuitively. Thus,…

音频与语音处理 · 电气工程与系统科学 2023-11-21 Yuke Lin , Xiaoyi Qin , Huahua Cui , Zhenyi Zhu , Ming Li

The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion…

音频与语音处理 · 电气工程与系统科学 2024-12-31 Ashishkumar Gudmalwar , Ishan D. Biyani , Nirmesh Shah , Pankaj Wasnik , Rajiv Ratn Shah

As the phonetic and acoustic manifestations of laughter in conversation are highly diverse, laughter synthesis should be capable of accommodating such diversity while maintaining high controllability. This paper proposes a generative model…

音频与语音处理 · 电气工程与系统科学 2023-09-01 Hiroki Mori , Shunya Kimura

Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Soumya Dutta

We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding the formation of emotions in visually grounded conversations. The task involves three skills: (1) Dialog-based…

State-of-the-art speech synthesis models try to get as close as possible to the human voice. Hence, modelling emotions is an essential part of Text-To-Speech (TTS) research. In our work, we selected FastSpeech2 as the starting point and…

音频与语音处理 · 电气工程与系统科学 2023-07-04 Daria Diatlova , Vitaly Shutov

Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion…

计算与语言 · 计算机科学 2018-04-02 Egor Lakomkin , Cornelius Weber , Sven Magg , Stefan Wermter

Accent plays a significant role in speech communication, influencing one's capability to understand as well as conveying a person's identity. This paper introduces a novel and efficient framework for accented Text-to-Speech (TTS) synthesis…

音频与语音处理 · 电气工程与系统科学 2024-10-01 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

We present a generative adversarial network to synthesize 3D pose sequences of co-speech upper-body gestures with appropriate affective expressions. Our network consists of two components: a generator to synthesize gestures from a joint…

多媒体 · 计算机科学 2024-11-26 Uttaran Bhattacharya , Elizabeth Childs , Nicholas Rewkowski , Dinesh Manocha

Emotional Video Captioning (EVC) is an emerging task, which aims to describe factual content with the intrinsic emotions expressed in videos. Existing works perceive global emotional cues and then combine with video content to generate…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Weidong Chen , Cheng Ye , Zhendong Mao , Peipei Song , Xinyan Liu , Lei Zhang , Xiaojun Chang , Yongdong Zhang

People communicate using both speech and non-verbal signals such as gestures, face expression or body pose. Non-verbal signals impact the meaning of the spoken utterance in an abundance of ways. An absence of non-verbal signals impoverishes…

计算机视觉与模式识别 · 计算机科学 2021-08-24 Manuel Rebol , Christian Gütl , Krzysztof Pietroszek

The field of affective computing focuses on recognizing, interpreting, and responding to human emotions, and has broad applications across education, child development, and human health and wellness. However, developing affective computing…

人工智能 · 计算机科学 2025-05-01 Emily Zhou , Khushboo Khatri , Yixue Zhao , Bhaskar Krishnamachari

The importance of modeling speech articulation for high-quality audiovisual (AV) speech synthesis is widely acknowledged. Nevertheless, while state-of-the-art, data-driven approaches to facial animation can make use of sophisticated motion…

人机交互 · 计算机科学 2012-09-25 Ingmar Steiner , Korin Richmond , Slim Ouni

Facial expressions are a form of non-verbal communication that humans perform seamlessly for meaningful transfer of information. Most of the literature addresses the facial expression recognition aspect however, with the advent of…

计算机视觉与模式识别 · 计算机科学 2022-02-09 J. Rafid Siddiqui

We present the JVNV, a Japanese emotional speech corpus with verbal content and nonverbal vocalizations whose scripts are generated by a large-scale language model. Existing emotional speech corpora lack not only proper emotional scripts…