English
Related papers

Related papers: ES4R: Speech Encoding Based on Prepositive Affecti…

200 papers

This report presents the Empathetic Cascading Networks (ECN) framework, a multi-stage prompting method designed to enhance the empathetic and inclusive capabilities of large language models. ECN employs four stages: Perspective Adoption,…

Computation and Language · Computer Science 2026-02-20 Wangjiaxuan Xin

Empathetic response generation aims to comprehend the cognitive and emotional states in dialogue utterances and generate proper responses. Psychological theories posit that comprehending emotional and cognitive states necessitates…

Computation and Language · Computer Science 2024-06-04 Zhou Yang , Zhaochun Ren , Yufeng Wang , Chao Chen , Haizhou Sun , Xiaofei Zhu , Xiangwen Liao

Cross-lingual Speech Emotion Recognition (CLSER) aims to identify emotional states in unseen languages. However, existing methods heavily rely on the semantic synchrony of complete labels and static feature stability, hindering low-resource…

Sound · Computer Science 2026-04-10 Ya Zhao , Yinfeng Yu , Liejun Wang

Current approaches to empathetic response generation typically encode the entire dialogue history directly and put the output into a decoder to generate friendly feedback. These methods focus on modelling contextual information but neglect…

Computation and Language · Computer Science 2023-11-28 Guoqing Lv , Jiang Li , Xiaoping Wang , Zhigang Zeng

This paper introduces EmpathyEar, a pioneering open-source, avatar-based multimodal empathetic chatbot, to fill the gap in traditional text-only empathetic response generation (ERG) systems. Leveraging the advancements of a large language…

Multimedia · Computer Science 2024-06-24 Hao Fei , Han Zhang , Bin Wang , Lizi Liao , Qian Liu , Erik Cambria

Emotion plays a fundamental role in human interaction, and therefore systems capable of identifying emotions in speech are crucial in the context of human-computer interaction. Speech emotion recognition (SER) is a challenging problem,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Lucas Ueda , João Lima , Leonardo Marques , Paula Costa

Automatic speech recognition (ASR) outcomes serve as input for downstream tasks, substantially impacting the satisfaction level of end-users. Hence, the diagnosis and enhancement of the vulnerabilities present in the ASR model bear…

Computation and Language · Computer Science 2024-01-29 Seonmin Koo , Chanjun Park , Jinsung Kim , Jaehyung Seo , Sugyeong Eo , Hyeonseok Moon , Heuiseok Lim

We introduce the task of expressive speech retrieval, where the goal is to retrieve speech utterances spoken in a given style based on a natural language description of that style. While prior work has primarily focused on performing speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-18 Wonjune Kang , Deb Roy

The purpose of emotion recognition in conversation (ERC) is to identify the emotion category of an utterance based on contextual information. Previous ERC methods relied on simple connections for cross-modal fusion and ignored the…

Computation and Language · Computer Science 2024-05-29 Haoxiang Shi , Xulong Zhang , Ning Cheng , Yong Zhang , Jun Yu , Jing Xiao , Jianzong Wang

Emotion Recognition in Conversation~(ERC) across modalities is of vital importance for a variety of applications, including intelligent healthcare, artificial intelligence for conversation, and opinion mining over chat history. The crux of…

Computation and Language · Computer Science 2023-06-07 Xingwei Liang , You Zou , Ruifeng Xu

Emotion Recognition in Conversations (ERC) is essential for building empathetic human-machine systems. Existing studies on ERC primarily focus on summarizing the context information in a conversation, however, ignoring the differentiated…

Computation and Language · Computer Science 2020-10-16 Yuzhao Mao , Qi Sun , Guang Liu , Xiaojie Wang , Weiguo Gao , Xuan Li , Jianping Shen

Affective computing is very important in the relationship between man and machine. In this paper, a system for speech emotion recognition (SER) based on speech signal is proposed, which uses new techniques in different stages of processing.…

Sound · Computer Science 2021-11-16 Fatemeh Daneshfar , Seyed Jahanshah Kabudian

Emotion recognition from speech is a challenging task. Re-cent advances in deep learning have led bi-directional recur-rent neural network (Bi-RNN) and attention mechanism as astandard method for speech emotion recognition, extractingand…

Sound · Computer Science 2021-06-09 Zixuan Peng , Yu Lu , Shengfeng Pan , Yunfeng Liu

Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classification problem, real-world affective states are often…

Sound · Computer Science 2026-02-05 Hong Jia , Weibin Li , Jingyao Wu , Xiaofeng Yu , Yan Gao , Jintao Cheng , Xiaoyu Tang , Feng Xia , Ting Dang

Large language models have revolutionized sign language generation by automatically transforming text into high-quality sign language videos, providing accessible communication for the Deaf community. However, existing LLM-based approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yanchao Zhao , Jihao Zhu , Yu Liu , Weizhuo Chen , Yuling Yang , Kun Peng

Most end-to-end (E2E) speech recognition models are composed of encoder and decoder blocks that perform acoustic and language modeling functions. Pretrained large language models (LLMs) have the potential to improve the performance of E2E…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-04 Shaoshi Ling , Yuxuan Hu , Shuangbei Qian , Guoli Ye , Yao Qian , Yifan Gong , Ed Lin , Michael Zeng

Emotional support is a crucial ability for many conversation scenarios, including social interactions, mental health support, and customer service chats. Following reasonable procedures and using various support skills can help to…

Computation and Language · Computer Science 2021-06-03 Siyang Liu , Chujie Zheng , Orianna Demasi , Sahand Sabour , Yu Li , Zhou Yu , Yong Jiang , Minlie Huang

Expressing empathy is important in everyday conversations, and exploring how empathy arises is crucial in automatic response generation. Most previous approaches consider only a single factor that affects empathy. However, in practice,…

Computation and Language · Computer Science 2022-12-06 Yangbin Chen , Chunfeng Liang

End-to-end models have achieved impressive results on the task of automatic speech recognition (ASR). For low-resource ASR tasks, however, labeled data can hardly satisfy the demand of end-to-end models. Self-supervised acoustic…

Computation and Language · Computer Science 2021-05-12 Cheng Yi , Shiyu Zhou , Bo Xu

The echo state network (ESN) is a powerful and efficient tool for displaying dynamic data. However, many existing ESNs have limitations for properly modeling high-dimensional data. The most important limitation of these networks is the high…

Sound · Computer Science 2021-11-16 Fatemeh Daneshfar , Seyed Jahanshah Kabudian