English
Related papers

Related papers: MSF-SER: Enriching Acoustic Modeling with Multi-Gr…

200 papers

Speech emotion recognition is crucial in human-computer interaction, but extracting and using emotional cues from audio poses challenges. This paper introduces MFHCA, a novel method for Speech Emotion Recognition using Multi-Spatial Fusion…

Sound · Computer Science 2024-04-23 Xinxin Jiao , Liejun Wang , Yinfeng Yu

Automatic emotion recognition (AER) based on enriched multimodal inputs, including text, speech, and visual clues, is crucial in the development of emotionally intelligent machines. Although complex modality relationships have been proven…

Multimedia · Computer Science 2021-09-16 Shuyun Tang , Zhaojie Luo , Guoshun Nan , Yuichiro Yoshikawa , Ishiguro Hiroshi

Multimodal emotion recognition from speech is an important area in affective computing. Fusing multiple data modalities and learning representations with limited amounts of labeled data is a challenging task. In this paper, we explore the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-08 Shamane Siriwardhana , Andrew Reis , Rivindu Weerasekera , Suranga Nanayakkara

With the rapid rise of social media and Internet culture, memes have become a popular medium for expressing emotional tendencies. This has sparked growing interest in Meme Emotion Understanding (MEU), which aims to classify the emotional…

Computation and Language · Computer Science 2025-11-17 Yi Shi , Wenlong Meng , Zhenyuan Guo , Chengkun Wei , Wenzhi Chen

This paper presents our contributions to the Speech Emotion Recognition in Naturalistic Conditions (SERNC) Challenge, where we address categorical emotion recognition and emotional attribute prediction. To handle the complexities of natural…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-15 Hyo Jin Jon , Longbin Jin , Hyuntaek Jung , Hyunseo Kim , Donghun Min , Eun Yi Kim

New-age conversational agent systems perform both speech emotion recognition (SER) and automatic speech recognition (ASR) using two separate and often independent approaches for real-world application in noisy environments. In this paper,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-29 Lokesh Bansal , S. Pavankumar Dubagunta , Malolan Chetlur , Pushpak Jagtap , Aravind Ganapathiraju

Multimodal emotion recognition (MER) is a fundamental complex research problem due to the uncertainty of human emotional expression and the heterogeneity gap between different modalities. Audio and text modalities are particularly important…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-07 Jiachen Luo , Huy Phan , Joshua Reiss

Speech emotion recognition (SER) is the task of recognising human's emotional states from speech. SER is extremely prevalent in helping dialogue systems to truly understand our emotions and become a trustworthy human conversational partner.…

Sound · Computer Science 2022-10-27 Zhao Ren , Thanh Tam Nguyen , Yi Chang , Björn W. Schuller

Multimodal emotion analysis is shifting from static classification to generative reasoning. Beyond simple label prediction, robust affective reasoning must synthesize fine-grained signals such as facial micro-expressions and prosodic which…

Multimedia · Computer Science 2026-02-05 Zhixian Zhao , Wenjie Tian , Lei Xie

Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Soumya Dutta

Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguistic information loss during quantization. This paper presents…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Esther Sun , Abinay Reddy Naini , Carlos Busso

Speech emotion recognition (SER) is vital for obtaining emotional intelligence and understanding the contextual meaning of speech. Variations of consonant-vowel (CV) phonemic boundaries can enrich acoustic context with linguistic cues,…

Sound · Computer Science 2023-07-03 Anna Ollerenshaw , Md Asif Jalal , Rosanna Milner , Thomas Hain

Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions conveyed through speech. While early SER systems relied on limited…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-25 Luz Martinez-Lucas , Pravin Mote , Abinay Reddy Naini , Mohammed Abdelwahab , Carlos Busso

We present a new end-to-end architecture for automatic speech recognition (ASR) that can be trained using \emph{symbolic} input in addition to the traditional acoustic input. This architecture utilizes two separate encoders: one for…

Computation and Language · Computer Science 2018-06-19 Adithya Renduchintala , Shuoyang Ding , Matthew Wiesner , Shinji Watanabe

Multimodal Emotion Recognition in Conversations (MERC) aims to classify utterance emotions using textual, auditory, and visual modal features. Most existing MERC methods assume each utterance has complete modalities, overlooking the common…

Computation and Language · Computer Science 2024-12-02 Fangze Fu , Wei Ai , Fan Yang , Yuntao Shou , Tao Meng , Keqin Li

Vision is often used as a complementary modality for audio speech recognition (ASR), especially in the noisy environment where performance of solo audio modality significantly deteriorates. After combining visual modality, ASR is upgraded…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Bo Xu , Cheng Lu , Yandong Guo , Jacob Wang

Speech Self-Supervised Learning (SSL) has demonstrated considerable efficacy in various downstream tasks. Nevertheless, prevailing self-supervised models often overlook the incorporation of emotion-related prior information, thereby…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Rui Liu , Zening Ma

Emotion estimation in music listening is confronting challenges to capture the emotion variation of listeners. Recent years have witnessed attempts to exploit multimodality fusing information from musical contents and physiological signals…

Artificial Intelligence · Computer Science 2016-12-01 Nattapong Thammasan , Ken-ichi Fukui , Masayuki Numao

The integration of information across multiple modalities and across time is a promising way to enhance the emotion recognition performance of affective systems. Much previous work has focused on instantaneous emotion recognition. The 2018…

Image and Video Processing · Electrical Eng. & Systems 2018-05-07 Didan Deng , Yuqian Zhou , Jimin Pi , Bertram E. Shi

We introduce PGF-Net (Progressive Gated-Fusion Network), a novel deep learning framework designed for efficient and interpretable multimodal sentiment analysis. Our framework incorporates three primary innovations. Firstly, we propose a…

Machine Learning · Computer Science 2025-08-25 Bin Wen , Tien-Ping Tan