English
Related papers

Related papers: Semantic Matters: Multimodal Features for Affectiv…

200 papers

Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable of comprehending diverse audio signal, performing audio analysis and generating textual responses. However, in speech emotion recognition (SER), ALMs often suffer…

Sound · Computer Science 2025-12-30 Zhixian Zhao , Xinfa Zhu , Xinsheng Wang , Shuiyuan Wang , Xuelong Geng , Wenjie Tian , Lei Xie

Multimodal learning pipelines have benefited from the success of pretrained language models. However, this comes at the cost of increased model parameters. In this work, we propose Adapted Multimodal BERT (AMB), a BERT-based architecture…

Computation and Language · Computer Science 2022-12-02 Odysseas S. Chlapanis , Georgios Paraskevopoulos , Alexandros Potamianos

The fifth Affective Behavior Analysis in-the-wild (ABAW) competition has multiple challenges such as Valence-Arousal Estimation Challenge, Expression Classification Challenge, Action Unit Detection Challenge, Emotional Reaction Intensity…

Computer Vision and Pattern Recognition · Computer Science 2023-03-20 Darshan Gera , Badveeti Naveen Siva Kumar , Bobbili Veerendra Raj Kumar , S Balasubramanian

With the rapid rise of social media and Internet culture, memes have become a popular medium for expressing emotional tendencies. This has sparked growing interest in Meme Emotion Understanding (MEU), which aims to classify the emotional…

Computation and Language · Computer Science 2025-11-17 Yi Shi , Wenlong Meng , Zhenyuan Guo , Chengkun Wei , Wenzhi Chen

Information on social media comprises of various modalities such as textual, visual and audio. NLP and Computer Vision communities often leverage only one prominent modality in isolation to study social media. However, the computational…

Computer Vision and Pattern Recognition · Computer Science 2020-08-11 Chhavi Sharma , Deepesh Bhageria , William Scott , Srinivas PYKL , Amitava Das , Tanmoy Chakraborty , Viswanath Pulabaigari , Bjorn Gamback

Visual-textual sentiment analysis aims to predict sentiment with the input of a pair of image and text, which poses a challenge in learning effective features for diverse input images. To address this, we propose a holistic method that…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Junyu Chen , Jie An , Hanjia Lyu , Christopher Kanan , Jiebo Luo

We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions. To support this framework, we expand the…

Sound · Computer Science 2025-11-21 Suhwan Choi , Kyu Won Kim , Myungjoo Kang

Multimodal aspect-based sentiment classification (MASC) is an emerging task due to an increase in user-generated multimodal content on social platforms, aimed at predicting sentiment polarity toward specific aspect targets (i.e., entities…

Computation and Language · Computer Science 2025-04-23 Luwei Xiao , Rui Mao , Shuai Zhao , Qika Lin , Yanhao Jia , Liang He , Erik Cambria

While Wav2Vec 2.0 has been proposed for speech recognition (ASR), it can also be used for speech emotion recognition (SER); its performance can be significantly improved using different fine-tuning strategies. Two baseline methods, vanilla…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-22 Li-Wei Chen , Alexander Rudnicky

This paper presents our system developed for SemEval-2026 Task 2. The task requires modeling both current affect and short-term affective change in chronologically ordered user-generated texts. We explore three complementary approaches: (1)…

Computation and Language · Computer Science 2026-05-28 Darya Hryhoryeva , Amaia Zurinaga , Hamidreza Jamalabadi , Iryna Gurevych

Speech emotion recognition is a crucial problem manifesting in a multitude of applications such as human computer interaction and education. Although several advancements have been made in the recent years, especially with the advent of…

Sound · Computer Science 2021-03-05 Panagiotis Tzirakis , Anh Nguyen , Stefanos Zafeiriou , Björn W. Schuller

Automatic emotion recognition is one of the central concerns of the Human-Computer Interaction field as it can bridge the gap between humans and machines. Current works train deep learning models on low-level data representations to solve…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-22 Mariana Rodrigues Makiuchi , Kuniaki Uto , Koichi Shinoda

The ACII Affective Vocal Bursts (A-VB) competition introduces a new topic in affective computing, which is understanding emotional expression using the non-verbal sound of humans. We are familiar with emotion recognition via verbal vocal or…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-04 Dang-Khanh Nguyen , Sudarshan Pant , Ngoc-Huynh Ho , Guee-Sang Lee , Soo-Huyng Kim , Hyung-Jeong Yang

This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Puneet Kumar , Sarthak Malik , Balasubramanian Raman

In recent years, deep learning has achieved innovative advancements in various fields, including the analysis of human emotions and behaviors. Initiatives such as the Affective Behavior Analysis in-the-wild (ABAW) competition have been…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Seongjae Min , Junseok Yang , Sangjun Lim , Junyong Lee , Sangwon Lee , Sejoon Lim

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the INTERSPEECH 2025…

Continuous affect prediction in the wild is a very interesting problem and is challenging as continuous prediction involves heavy computation. This paper presents the methodologies and techniques used in our contribution to predict…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-02 Sowmya Rasipuram , Junaid Hamid Bhat , Anutosh Maitra

This project intends to study the image representation based on attention mechanism and multimodal data. By adding multiple pattern layers to the attribute model, the semantic and hidden layers of image content are integrated. The word…

Computation and Language · Computer Science 2024-06-14 Dan Sun , Yaxin Liang , Yining Yang , Yuhan Ma , Qishi Zhan , Erdi Gao

The ability to understand emotions is an essential component of human-like artificial intelligence, as emotions greatly influence human cognition, decision making, and social interactions. In addition to emotion recognition in…

Computation and Language · Computer Science 2024-07-09 Fanfan Wang , Heqing Ma , Jianfei Yu , Rui Xia , Erik Cambria

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech…

Computation and Language · Computer Science 2020-04-06 Haiyang Xu , Hui Zhang , Kun Han , Yun Wang , Yiping Peng , Xiangang Li