English
Related papers

Related papers: Calibrating Multimodal Consensus for Emotion Recog…

200 papers

Multimodal music emotion recognition (MMER) is an emerging discipline in music information retrieval that has experienced a surge in interest in recent years. This survey provides a comprehensive overview of the current state-of-the-art in…

Multimedia · Computer Science 2025-04-29 Rashini Liyanarachchi , Aditya Joshi , Erik Meijering

In this work, we present a novel method for music emotion recognition that leverages Large Language Model (LLM) embeddings for label alignment across multiple datasets and zero-shot prediction on novel categories. First, we compute LLM…

Sound · Computer Science 2024-10-18 Renhang Liu , Abhinaba Roy , Dorien Herremans

Compound Expression Recognition (CER), a subfield of affective computing, aims to detect complex emotional states formed by combinations of basic emotions. In this work, we present a novel zero-shot multimodal approach for CER that combines…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Elena Ryumina , Maxim Markitantov , Alexandr Axyonov , Dmitry Ryumin , Mikhail Dolgushin , Alexey Karpov

Emotion recognition is an important research direction in artificial intelligence, helping machines understand and adapt to human emotional states. Multimodal electrophysiological(ME) signals, such as EEG, GSR, respiration(Resp), and…

Multimedia · Computer Science 2023-08-07 Yunfei Guo , Tao Zhang , Wu Huang

Recent advancements in Unified Multimodal Models (UMMs) have enabled remarkable image understanding and generation capabilities. However, while models like Gemini-2.5-Flash-Image show emerging abilities to reason over multiple related…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Mingrui Wu , Hang Liu , Jiayi Ji , Xiaoshuai Sun , Rongrong Ji

This paper presents InterMPL, a semi-supervised learning method of end-to-end automatic speech recognition (ASR) that performs pseudo-labeling (PL) with intermediate supervision. Momentum PL (MPL) trains a connectionist temporal…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-20 Yosuke Higuchi , Tetsuji Ogawa , Tetsunori Kobayashi , Shinji Watanabe

Model calibration seeks to ensure that models produce confidence scores that accurately reflect the true likelihood of their predictions being correct. However, existing calibration approaches are fundamentally tied to datasets of one-hot…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Haoyang Luo , Linwei Tao , Minjing Dong , Chang Xu

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive…

In recent years, speech emotion recognition technology is of great significance in industrial applications such as call centers, social robots and health care. The combination of speech recognition and speech emotion recognition can improve…

Artificial Intelligence · Computer Science 2023-02-20 Ying Zhou , Xuefeng Liang , Yu Gu , Yifei Yin , Longshan Yao

This study investigates the integration of trustworthy prior reasoning knowledge from MLLMs into multimodal emotion recognition. We employ Gemini to generate fine-grained, modality-separable reasoning traces, which are injected as priors…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Zhepeng Wang , Yingjian Zhu , Guanghao Dong , Hongzhu Yi , Feng Chen , Xinming Wang , Jun Xie

The lack of data and the difficulty of multimodal fusion have always been challenges for multimodal emotion recognition (MER). In this paper, we propose to use pretrained models as upstream network, wav2vec 2.0 for audio modality and BERT…

Computation and Language · Computer Science 2023-02-28 Dekai Sun , Yancheng He , Jiqing Han

In terms of human-computer interaction, it is becoming more and more important to correctly understand the user's emotional state in a conversation, so the task of multimodal emotion recognition (MER) started to receive more attention.…

Computation and Language · Computer Science 2024-01-04 Wei Ai , FuChen Zhang , Tao Meng , YunTao Shou , HongEn Shao , Keqin Li

Multimodal emotion recognition in conversations (MERC) requires integrating multimodal signals while being robust to noise and modeling contextual reasoning. Existing approaches often emphasize fusion but overlook uncertainty in noisy…

Computation and Language · Computer Science 2026-04-03 Yiqiang Cai , Chengyan Wu , Bolei Ma , Bo Chen , Yun Xue , Julia Hirschberg , Ziwei Gong

Concept Bottleneck Models (CBMs) assume that training examples (e.g., x-ray images) are annotated with high-level concepts (e.g., types of abnormalities), and perform classification by first predicting the concepts, followed by predicting…

Computation and Language · Computer Science 2023-12-19 Danis Alukaev , Semen Kiselev , Ilya Pershin , Bulat Ibragimov , Vladimir Ivanov , Alexey Kornaev , Ivan Titov

Multimodal semantic communication has gained widespread attention due to its ability to enhance downstream task performance. A key challenge in such systems is the effective fusion of features from different modalities, which requires the…

Image and Video Processing · Electrical Eng. & Systems 2025-09-03 Haoshuo Zhang , Yufei Bo , Hongwei Zhang , Meixia Tao

Emotion recognition from multi-modal physiological and behavioral signals plays a pivotal role in affective computing, yet most existing models remain constrained to the prediction of singular emotions in controlled laboratory settings.…

Machine Learning · Computer Science 2026-02-25 Ming Li , Yong-Jin Liu , Fang Liu , Huankun Sheng , Yeying Fan , Yixiang Wei , Minnan Luo , Weizhan Zhang , Wenping Wang

Multi-modal emotion recognition has garnered increasing attention as it plays a significant role in human-computer interaction (HCI) in recent years. Since different discrete emotions may exist at the same time, compared with single-class…

Machine Learning · Computer Science 2025-07-29 Chuhang Zheng , Chunwei Tian , Jie Wen , Daoqiang Zhang , Qi Zhu

Multimodal speech emotion recognition aims to detect speakers' emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while…

Computation and Language · Computer Science 2023-04-11 Zhen Wu , Yizhe Lu , Xinyu Dai

Multi-label classification (MLC) faces challenges from label noise in training data due to annotating diverse semantic labels for each image. Current methods mainly target identifying and correcting label mistakes using trained MLC models,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Zhixiang Yuan , Kaixin Zhang , Tao Huang

Human multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches have achieved promising results, they still struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Yong Li , Yuanzhi Wang , Yi Ding , Shiqing Zhang , Ke Lu , Cuntai Guan