English
Related papers

Related papers: A Study on the Data Distribution Gap in Music Emot…

200 papers

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio…

Sound · Computer Science 2025-07-30 Ohad Cohen , Gershon Hazan , Sharon Gannot

Collecting large, aligned cross-modal datasets for music-flavor research is difficult because perceptual experiments are costly and small by design. We address this bottleneck through two complementary experiments. The first tests whether…

Sound · Computer Science 2026-04-14 Matteo Spanio , Valentina Frezzato , Antonio Rodà

Musical features and descriptors could be coarsely divided into three levels of complexity. The bottom level contains the basic building blocks of music, e.g., chords, beats and timbre. The middle level contains concepts that emerge from…

Sound · Computer Science 2018-06-14 Anna Aljanaki , Mohammad Soleymani

Emotions manifest through physical experiences and bodily reactions, yet identifying such embodied emotions in text remains understudied. We present an embodied emotion classification dataset, CHEER-Ekman, extending the existing binary…

Computation and Language · Computer Science 2025-09-26 Phan Anh Duong , Cat Luong , Divyesh Bommana , Tianyu Jiang

In this work, we present a lightweight and privacy-preserving Multimodal Emotion Recognition (MER) framework designed for deployment on edge devices. To demonstrate framework's versatility, our implementation uses three modalities - speech,…

Artificial Intelligence · Computer Science 2026-02-11 Rémi Grzeczkowicz , Eric Soriano , Ali Janati , Miyu Zhang , Gerard Comas-Quiles , Victor Carballo Araruna , Aneesh Jonelagadda

Music evokes profound emotions, yet the universality of emotional descriptors across languages remains debated. A key challenge in cross-cultural research on music emotion is biased stimulus selection and manual curation of taxonomies,…

Computation and Language · Computer Science 2025-02-14 Elif Celen , Pol van Rijn , Harin Lee , Nori Jacoby

The subjective perception of emotion leads to inconsistent labels from human annotators. Typically, utterances lacking majority-agreed labels are excluded when training an emotion classifier, which cause problems when encountering ambiguous…

Computation and Language · Computer Science 2024-10-14 Wen Wu , Bo Li , Chao Zhang , Chung-Cheng Chiu , Qiujia Li , Junwen Bai , Tara N. Sainath , Philip C. Woodland

In automatic emotion recognition (AER), labels assigned by different human annotators to the same utterance are often inconsistent due to the inherent complexity of emotion and the subjectivity of perception. Though deterministic labels…

Sound · Computer Science 2024-04-02 Wen Wu , Chao Zhang , Philip C. Woodland

Recognizing emotions from text in multimodal architectures has yielded promising results, surpassing video and audio modalities under certain circumstances. However, the method by which multimodal data is collected can be significant for…

Machine Learning · Computer Science 2021-03-08 A. Sutherland , S. Magg , C. Weber , S. Wermter

Emotion Recognition (ER) is the process of identifying human emotions from given data. Currently, the field heavily relies on facial expression recognition (FER) because facial expressions contain rich emotional cues. However, it is…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Yuanyuan Liu , Lin Wei , Kejun Liu , Yibing Zhan , Zijing Chen , Zhe Chen , Shiguang Shan

Emotion recognition algorithms rely on data annotated with high quality labels. However, emotion expression and perception are inherently subjective. There is generally not a single annotation that can be unambiguously declared "correct".…

Multimodal Emotion Recognition (MER) aims to perceive human emotions through three modes: language, vision, and audio. Previous methods primarily focused on modal fusion without adequately addressing significant distributional differences…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Jichao Zhu , Jun Yu

This work was developed aiming to employ Statistical techniques to the field of Music Emotion Recognition, a well-recognized area within the Signal Processing world, but hardly explored from the statistical point of view. Here, we opened…

Machine Learning · Statistics 2021-07-13 Nathalie Deziderio , Hugo Tremonte de Carvalho

We introduce the problem of learning affective correspondence between audio (music) and visual data (images). For this task, a music clip and an image are considered similar (having true correspondence) if they have similar emotion content.…

Multimedia · Computer Science 2019-04-18 Gaurav Verma , Eeshan Gunesh Dhekane , Tanaya Guha

Multimodal Emotion Recognition (MER) aims to accurately identify human emotional states by integrating heterogeneous modalities such as visual, auditory, and textual data. Existing approaches predominantly rely on unified emotion labels to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Wen Yin , Siyu Zhan , Cencen Liu , Xin Hu , Guiduo Duan , Xiurui Xie , Yuan-Fang Li , Tao He

Speech emotion recognition (SER) systems often exhibit gender bias. However, the effectiveness and robustness of existing debiasing methods in such multi-label scenarios remain underexplored. To address this gap, we present EMO-Debias, a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-06 Yi-Cheng Lin , Huang-Cheng Chou , Yu-Hsuan Li Liang , Hung-yi Lee

In affective computing, datasets often contain multiple annotations from different annotators, which may lack full agreement. Typically, these annotations are merged into a single gold standard label, potentially losing valuable inter-rater…

Human-Computer Interaction · Computer Science 2025-05-28 Ibrahim Shoer , Engin Erzin

Music emotion recognition (MER), a sub-task of music information retrieval (MIR), has developed rapidly in recent years. However, the learning of affect-salient features remains a challenge. In this paper, we propose an end-to-end…

Sound · Computer Science 2022-07-01 Zi Huang , Shulei Ji , Zhilan Hu , Chuangjian Cai , Jing Luo , Xinyu Yang

Multi-modal conversation emotion recognition (MCER) aims to recognize and track the speaker's emotional state using text, speech, and visual information in the conversation scene. Analyzing and studying MCER issues is significant to…

Artificial Intelligence · Computer Science 2025-11-14 Yuntao Shou , Tao Meng , Wei Ai , Fangze Fu , Nan Yin , Keqin Li

Whether literally or suggestively, the concept of soundscape is alluded in both modern and ancient music. In this study, we examine whether we can analyze and compare Western and Chinese classical music based on soundscape models. We…

Sound · Computer Science 2020-02-24 Jianyu Fan , Yi-Hsuan Yang , Kui Dong , Philippe Pasquier
‹ Prev 1 3 4 5 6 7 10 Next ›