English
Related papers

Related papers: Creating Emordle: Animating Word Cloud for Emotion…

200 papers

This paper presents Project Riley, a novel multimodal and multi-model conversational AI architecture oriented towards the simulation of reasoning influenced by emotional states. Drawing inspiration from Pixar's Inside Out, the system…

Artificial Intelligence · Computer Science 2025-09-09 Ana Rita Ortigoso , Gabriel Vieira , Daniel Fuentes , Luis Frazão , Nuno Costa , António Pereira

The task of audio-driven portrait animation involves generating a talking head video using an identity image and an audio track of speech. While many existing approaches focus on lip synchronization and video quality, few tackle the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Jian Zhang , Weijian Mai , Zhijun Zhang

Automated emotion detection in speech is a challenging task due to the complex interdependence between words and the manner in which they are spoken. It is made more difficult by the available datasets; their small size and incompatible…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-16 Amith Ananthram , Kailash Karthik Saravanakumar , Jessica Huynh , Homayoon Beigi

Humans' experience of the world is profoundly multimodal from the beginning, so why do existing state-of-the-art language models only use text as a modality to learn and represent semantic meaning? In this paper we review the literature on…

Computation and Language · Computer Science 2021-05-12 Casey Kennington

In this paper, we address three challenges in utterance-level emotion recognition in dialogue systems: (1) the same word can deliver different emotions in different contexts; (2) some emotions are rarely seen in general dialogues; (3)…

Computation and Language · Computer Science 2019-04-10 Wenxiang Jiao , Haiqin Yang , Irwin King , Michael R. Lyu

Explicitly modeling emotions in dialogue generation has important applications, such as building empathetic personal companions. In this study, we consider the task of expressing a specific emotion for dialogue generation. Previous…

Computation and Language · Computer Science 2021-09-23 Chengzhang Dong , Chenyang Huang , Osmar Zaïane , Lili Mou

Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-07 Tianhua Qi , Wenming Zheng , Björn W. Schuller , Zhaojie Luo , Haizhou Li

Dynamic facial expression recognition (FER) databases provide important data support for affective computing and applications. However, most FER databases are annotated with several basic mutually exclusive emotional categories and contain…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Yuanyuan Liu , Wei Dai , Chuanxu Feng , Wenbin Wang , Guanghao Yin , Jiabei Zeng , Shiguang Shan

Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predefined labels, reference audio, or natural language prompts.…

Computation and Language · Computer Science 2025-09-26 Sirui Wang , Andong Chen , Tiejun Zhao

Emotion recognition has a pivotal role in affective computing and in human-computer interaction. The current technological developments lead to increased possibilities of collecting data about the emotional state of a person. In general,…

Computer Vision and Pattern Recognition · Computer Science 2020-07-10 Andreea Birhala , Catalin Nicolae Ristea , Anamaria Radoi , Liviu Cristian Dutu

The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-31 Ashishkumar Gudmalwar , Ishan D. Biyani , Nirmesh Shah , Pankaj Wasnik , Rajiv Ratn Shah

In this paper, we propose a multimodal framework for speech emotion recognition that leverages entropy-aware score selection to combine speech and textual predictions. The proposed method integrates a primary pipeline that consists of an…

Sound · Computer Science 2025-08-29 ChenYi Chua , JunKai Wong , Chengxin Chen , Xiaoxiao Miao

Emotion recognition and generation have emerged as crucial topics in Artificial Intelligence research, playing a significant role in enhancing human-computer interaction within healthcare, customer service, and other fields. Although…

Machine Learning · Computer Science 2025-02-12 Rebecca Mobbs , Dimitrios Makris , Vasileios Argyriou

Multimodal speech emotion recognition aims to detect speakers' emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while…

Computation and Language · Computer Science 2023-04-11 Zhen Wu , Yizhe Lu , Xinyu Dai

Previous research on empathetic dialogue systems has mostly focused on generating responses given certain emotions. However, being empathetic not only requires the ability of generating emotional responses, but more importantly, requires…

Computation and Language · Computer Science 2019-08-22 Zhaojiang Lin , Andrea Madotto , Jamin Shin , Peng Xu , Pascale Fung

Emotion recognition and classification is a very active area of research. In this paper, we present a first approach to emotion classification using persistent entropy and support vector machines. A topology-based model is applied to obtain…

Sound · Computer Science 2019-03-22 R. Gonzalez-Diaz , E. Paluzo-Hidalgo , J. F. Quesada

Emotion recognition plays a pivotal role in enhancing human-computer interaction, particularly in movie recommendation systems where understanding emotional content is essential. While multimodal approaches combining audio and video have…

Sound · Computer Science 2025-11-25 Xiangrui Xiong , Zhou Zhou , Guocai Nong , Junlin Deng , Ning Wu

Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not…

Computation and Language · Computer Science 2023-12-20 Rui Liu , Yifan Hu , Yi Ren , Xiang Yin , Haizhou Li

Emotion Recognition in Conversation (ERC) involves detecting the underlying emotion behind each utterance within a conversation. Effectively generating representations for utterances remains a significant challenge in this task. Recent…

Computation and Language · Computer Science 2024-04-01 Fangxu Yu , Junjie Guo , Zhen Wu , Xinyu Dai

Speech emotion recognition is a challenging task because the emotion expression is complex, multimodal and fine-grained. In this paper, we propose a novel multimodal deep learning approach to perform fine-grained emotion recognition from…

Sound · Computer Science 2021-07-16 Hang Li , Wenbiao Ding , Zhongqin Wu , Zitao Liu