English
Related papers

Related papers: Towards Multimodal Empathetic Response Generation:…

200 papers

Multimodal Emotion Recognition (MER) aims to perceive human emotions through three modes: language, vision, and audio. Previous methods primarily focused on modal fusion without adequately addressing significant distributional differences…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Jichao Zhu , Jun Yu

Empathy is critical to successful mental health support. Empathy measurement has predominantly occurred in synchronous, face-to-face settings, and may not translate to asynchronous, text-based contexts. Because millions of people use…

Computation and Language · Computer Science 2020-09-18 Ashish Sharma , Adam S. Miner , David C. Atkins , Tim Althoff

As LLMs exhibit a high degree of human-like capability, increasing attention has been paid to role-playing research areas in which responses generated by LLMs are expected to mimic human replies. This has promoted the exploration of…

Artificial Intelligence · Computer Science 2024-10-31 Le Huang , Hengzhi Lan , Zijun Sun , Chuan Shi , Ting Bai

Retrieval-Augmented Generation (RAG) has become a core paradigm in document question answering tasks. However, existing methods have limitations when dealing with multimodal documents: one category of methods relies on layout analysis and…

Computation and Language · Computer Science 2026-03-09 Wang Chen , Wenhan Yu , Guanqiang Qi , Weikang Li , Yang Li , Lei Sha , Deguo Xia , Jizhou Huang

In this work, we present a lightweight and privacy-preserving Multimodal Emotion Recognition (MER) framework designed for deployment on edge devices. To demonstrate framework's versatility, our implementation uses three modalities - speech,…

Artificial Intelligence · Computer Science 2026-02-11 Rémi Grzeczkowicz , Eric Soriano , Ali Janati , Miyu Zhang , Gerard Comas-Quiles , Victor Carballo Araruna , Aneesh Jonelagadda

One challenge for dialogue agents is recognizing feelings in the conversation partner and replying accordingly, a key communicative skill. While it is straightforward for humans to recognize and acknowledge others' feelings in a…

Computation and Language · Computer Science 2019-08-30 Hannah Rashkin , Eric Michael Smith , Margaret Li , Y-Lan Boureau

In the literature, existing human-centric emotional motion generation methods primarily focus on boosting performance within a single scale-fixed dataset, largely neglecting the flexible and scale-increasing motion scenarios (e.g., sports,…

Artificial Intelligence · Computer Science 2025-12-23 Jiawen Wang , Jingjing Wang Tianyang Chen , Min Zhang , Guodong Zhou

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Affective judgment in real interaction is rarely a purely local prediction problem. Emotional meaning often depends on prior trajectory, accumulated context, and multimodal evidence that may be weak, noisy, or incomplete at the current…

Artificial Intelligence · Computer Science 2026-03-25 Deliang Wen , Ke Sun , Yu Wang

Multimodal Emotion Recognition in Conversation (MERC) aims to predict speakers' emotions by integrating textual, acoustic, and visual cues. Existing approaches either struggle to capture complex cross-modal interactions or experience…

Multimedia · Computer Science 2026-03-24 Xiaosen Lyu , Jiayu Xiong , Yuren Chen , Wanlong Wang , Xiaoqing Dai , Jing Wang

In affective computing, the task of Emotion Recognition in Conversations (ERC) has emerged as a focal area of research. The primary objective of this task is to predict emotional states within conversations by analyzing multimodal data…

Multimedia · Computer Science 2024-11-22 Xiaomin Yu , Feiyang Wang , Ziyue Qiao

Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving the recognition accuracy. However, the lack of high-quality…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-01 Jinming Chen , Jingyi Fang , Yuanzhong Zheng , Yaoxuan Wang , Haojun Fei

Empathetic response generation is increasingly significant in AI, necessitating nuanced emotional and cognitive understanding coupled with articulate response expression. Current large language models (LLMs) excel in response expression;…

Human-Computer Interaction · Computer Science 2024-02-20 Zhou Yang , Zhaochun Ren , Wang Yufeng , Shizhong Peng , Haizhou Sun , Xiaofei Zhu , Xiangwen Liao

Large language models (LLMs) have made significant progress in Emotional Intelligence (EI) and long-context modeling. However, existing benchmarks often overlook the fact that emotional information processing unfolds as a continuous…

Computation and Language · Computer Science 2026-01-13 Weichu Liu , Jing Xiong , Yuxuan Hu , Zixuan Li , Minghuan Tan , Ningning Mao , Hui Shen , Wendong Xu , Chaofan Tao , Min Yang , Chengming Li , Lingpeng Kong , Ngai Wong

Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large language models excel at multimodal reasoning, they typically treat emotion categories as…

Machine Learning · Computer Science 2026-05-20 Zeheng Wang , Bo Zhao , Yijie Zhu , Zhishu Liu , Hui Ma , Ruixin Zhang , Shouhong Ding , Qianyu Xie , Zitong Yu

The integration of conversational agents into our daily lives has become increasingly common, yet many of these agents cannot engage in deep interactions with humans. Despite this, there is a noticeable shortage of datasets that capture…

Human-Computer Interaction · Computer Science 2025-03-19 Mohammed Althubyani , Zhijin Meng , Shengyuan Xie , Cha Seung , Imran Razzak , Eduardo B. Sandoval , Baki Kocaballi , Francisco Cruz

In natural face-to-face interaction, participants seamlessly alternate between speaking and listening, producing facial behaviors (FBs) that are finely informed by long-range context and naturally exhibit contextual appropriateness and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Xiangyu Kong , Xiaoyu Jin , Yihan Pan , Haoqin Sun , Hengde Zhu , Xiaoming Xu , Xiaoming Wei , Lu Liu , Siyang Song

Talking Head Generation (THG) has emerged as a transformative technology in computer vision, enabling the synthesis of realistic human faces synchronized with image, audio, text, or video inputs. This paper provides a comprehensive review…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Vineet Kumar Rakesh , Soumya Mazumdar , Research Pratim Maity , Sarbajit Pal , Amitabha Das , Tapas Samanta

This work focuses on full-body co-speech gesture generation. Existing methods typically employ an autoregressive model accompanied by vector-quantized tokens for gesture generation, which results in information loss and compromises the…

Graphics · Computer Science 2025-03-19 Binjie Liu , Lina Liu , Sanyi Zhang , Songen Gu , Yihao Zhi , Tianyi Zhu , Lei Yang , Long Ye

Researches on dialogue empathy aim to endow an agent with the capacity of accurate understanding and proper responding for emotions. Existing models for empathetic dialogue generation focus on the emotion flow in one direction, that is,…

Computation and Language · Computer Science 2021-09-21 Lei Shen , Jinchao Zhang , Jiao Ou , Xiaofang Zhao , Jie Zhou
‹ Prev 1 4 5 6 7 8 10 Next ›